perf(traverse): reduce string operations in get_var_name_from_node#24007
Conversation
How to use the Graphite Merge QueueAdd either label to this PR to merge it via the merge queue:
You must have a Graphite account in order to use the merge queue. Sign up using this link. An organization admin has enabled the Graphite Merge Queue in this repository. Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue. This stack of pull requests is managed by Graphite. Learn more about stacking. |
Merging this PR will not alter performance
Comparing Footnotes
|
There was a problem hiding this comment.
Pull request overview
This PR optimizes oxc_traverse’s get_var_name_from_node by avoiding unnecessary string growth/truncation work and leveraging StringExt’s unchecked push APIs under a strict max-length invariant, while also correcting UTF-8 truncation behavior for Unicode boundary cases.
Changes:
- Stop concatenating once the variable name reaches
MAX_LEN(20 bytes) to avoid wasted work and allocations. - Use
StringExt::{push_unchecked, push_str_unchecked}to remove repeated “needs to grow?” checks in the hot path. - Expand and update unit tests to cover ASCII + multi-byte UTF-8 truncation and multi-part joining behavior.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
| crates/oxc_traverse/src/ast_operations/gather_node_parts.rs | Reworks get_var_name_from_node to cap work at 20 bytes, uses unchecked string extension methods, and adds extensive truncation/joining tests. |
| crates/oxc_traverse/Cargo.toml | Enables the string_ext feature on oxc_data_structures to use StringExt. |
38328f2 to
5ec98cf
Compare
d72530c to
2e62012
Compare
5ec98cf to
4d11f8b
Compare
4d11f8b to
47e000c
Compare
Merge activity
|
…24007) Improve the performance of creating a var name from an AST node. ## Perf 2 different optimizations: #### 1. Avoid building a huge string Previously we concatenated all identifiers into a string of potentially large length, then truncated it to 20 bytes at the end. Instead, once the string reaches the maximum length, stop adding to it. This avoids pointless work, and avoids extra allocations to grow the `String`. #### 2. Remove "needs to grow?" checks The first optimization ensures that the `String` never exceeds 20 bytes. Utilize this invariant to remove "does the `String` need to grow?" checks on every push to the `String`. This uses the `push_unchecked` and `push_str_unchecked` methods added to `String` in #24006. Adverserial review by Claude found the implementation to be sound, and it performed fuzzing using 500,000 random cases to test the implementation. Tests which are added in this PR also cover all the edge cases. ## Fix This PR also fixes incorrect truncation when some of the parts contain Unicode characters. Previously the string would be truncated to be shorter than 20 bytes when the last char before the cut point is a complete multi-byte (Unicode) character. Now it checks whether to truncate to less than 20 bytes based on whether the 21st byte is a character sequence start byte, instead of whether the 20th byte is a continuation character. The tests are altered to reflect this fix.
47e000c to
a55e0be
Compare
…24007) Improve the performance of creating a var name from an AST node. ## Perf 2 different optimizations: #### 1. Avoid building a huge string Previously we concatenated all identifiers into a string of potentially large length, then truncated it to 20 bytes at the end. Instead, once the string reaches the maximum length, stop adding to it. This avoids pointless work, and avoids extra allocations to grow the `String`. #### 2. Remove "needs to grow?" checks The first optimization ensures that the `String` never exceeds 20 bytes. Utilize this invariant to remove "does the `String` need to grow?" checks on every push to the `String`. This uses the `push_unchecked` and `push_str_unchecked` methods added to `String` in #24006. Adverserial review by Claude found the implementation to be sound, and it performed fuzzing using 500,000 random cases to test the implementation. Tests which are added in this PR also cover all the edge cases. ## Fix This PR also fixes incorrect truncation when some of the parts contain Unicode characters. Previously the string would be truncated to be shorter than 20 bytes when the last char before the cut point is a complete multi-byte (Unicode) character. Now it checks whether to truncate to less than 20 bytes based on whether the 21st byte is a character sequence start byte, instead of whether the 20th byte is a continuation character. The tests are altered to reflect this fix.
### 🚀 Features - 260425f semantic/examples: Include unresolved references (#24214) (camc314) - 2d9b0b3 minifier: Fold boolean-literal ternary branches in value contexts (#24110) (Dunqing) - 61fbf10 ast: Implement `ReplaceWith` on all AST types (#24013) (overlookmotel) - 7db7a29 allocator: Add `ReplaceWith` trait (#24012) (overlookmotel) - 4eb074e mangler: Add `reserved` option for names that must not be mangled (#24041) (Dunqing) - 2e62012 data_structures: Add `StringExt` trait (#24006) (overlookmotel) - 60e7160 minifier: Drop side-effect-free IIFEs whose result is unused (#23967) (Dunqing) - 26dd9e2 ast: Add method to widen inherited enum ref to parent ref (#23961) (overlookmotel) ### 🐛 Bug Fixes - e8b50ee transformer: Clean up semantics for stripped TypeScript syntax (#24180) (camc314) - d966d0b react_compiler: Remove clippy allows (#24168) (Boshen) - 854ef8d react_compiler: Compile generic functions instead of over-bailing on type-param hoisting (#24158) (Boshen) - 093586c react_compiler: Align memoization cache-slot allocation with Babel (#24157) (Boshen) - 09c8f59 react_compiler: Normalize snapshot fixture paths (#24142) (camc314) - f13df97 react_compiler: Drop stray empty statement from catch bindings (#24133) (Boshen) - cb2a505 react_compiler: Codegen destructuring reassignment targets (#24131) (Boshen) - b82c394 react_compiler: Propagate codegen invariants instead of emitting empty bodies (#24128) (Boshen) - 5771982 react_compiler: Render unchanged programs as source in fixture snapshots (#24129) (Boshen) - 4b16e1a transformer/async-to-generator: Preserve direct eval scope flags (#24136) (camc314) - 4e9194f react_compiler: Lower `delete obj.prop` to Property/ComputedDelete (#24123) (Boshen) - 0b25582 ast: Type binding node `typeAnnotation` as `TSTypeAnnotation | null` (#23113) (Boshen) - 018c0e5 transformer: Hoist lowered async declarations (#22770) (camc314) - 652fbaf mangler: Keep names of destructured exported bindings (#24036) (Dunqing) - e274415 minifier: Don't drop global calls that throw despite pure arguments (#23917) (Dunqing) - 59abb30 minifier: Only merge string literals in `try_fold_add` when the inner operator is `+` (#23622) (Jerry Zhao) ### ⚡ Performance - c5ca77b transformer: Avoid cloning refresh options (#24191) (camc314) - bf1a151 react_compiler: Compile out debug printers (#24184) (Boshen) - abb44a0 transformer: Build fixed object-rest arguments (#24190) (camc314) - a4db731 isolated_declarations: Use `ReplaceWith` instead of `TakeIn` (#24016) (overlookmotel) - ff10855 transformer: Use `ReplaceWith` instead of `TakeIn` (#24015) (overlookmotel) - bd49aff ecmascript: Avoid heap-allocating Math.min/max/imul operands (#23941) (Lawrence Lin) - e4b708b react_compiler: Skip compiled files before prefilters (#24171) (Boshen) - c59f2fe rust: Return impl ExactSizeIterator from slice-backed accessors (#24144) (Boshen) - 5d6d04a codegen: SWAR-skip boring byte runs in sourcemap line/column scan (#24023) (Boshen) - a55e0be traverse: Reduce string operations in `get_var_name_from_node` (#24007) (overlookmotel) - e6d48e1 transformer/nullish_coalescing: Move cold path into separate function (#23989) (overlookmotel) - c4e35b5 transformer/object_rest_spread: Pre-allocate capacity in `Vec` (#23988) (overlookmotel) - 527b8e5 transformer/decorators: Narrow type earlier (#23987) (overlookmotel) ### 📚 Documentation - 30d17f5 allocator: Clarify docs for `TakeIn::take_in_box` (#24093) (overlookmotel) - 675e6a8 ast: Correct doc comment for `PrivateFieldExpression` (#24008) (overlookmotel) - e4c30e6 minifier: Explain what `dce` mode means (#23994) (Dunqing) - 37cbf88 ast_macros: Document fields of `StructDetails` (#23959) (overlookmotel) - 4de3e54 ast: Correct doc comment (#23948) (overlookmotel) Co-authored-by: Boshen <1430279+Boshen@users.noreply.github.com>

Improve the performance of creating a var name from an AST node.
Perf
2 different optimizations:
1. Avoid building a huge string
Previously we concatenated all identifiers into a string of potentially large length, then truncated it to 20 bytes at the end.
Instead, once the string reaches the maximum length, stop adding to it. This avoids pointless work, and avoids extra allocations to grow the
String.2. Remove "needs to grow?" checks
The first optimization ensures that the
Stringnever exceeds 20 bytes. Utilize this invariant to remove "does theStringneed to grow?" checks on every push to theString.This uses the
push_uncheckedandpush_str_uncheckedmethods added toStringin #24006.Adverserial review by Claude found the implementation to be sound, and it performed fuzzing using 500,000 random cases to test the implementation. Tests which are added in this PR also cover all the edge cases.
Fix
This PR also fixes incorrect truncation when some of the parts contain Unicode characters.
Previously the string would be truncated to be shorter than 20 bytes when the last char before the cut point is a complete multi-byte (Unicode) character. Now it checks whether to truncate to less than 20 bytes based on whether the 21st byte is a character sequence start byte, instead of whether the 20th byte is a continuation character.
The tests are altered to reflect this fix.