Skip to content

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write - #65324

Closed
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding
Closed

src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write#65324
codebytere wants to merge 2 commits into
nodejs:mainfrom
codebytere:perf/src-simdutf-utf8-transcoding

Conversation

@codebytere

@codebytere codebytere commented Aug 16, 2026

Copy link
Copy Markdown
Member

StringDecoder gets 3–12× faster on UTF-8, and writing non-ASCII strings as UTF-8 (Buffer#write, Buffer.from(string), every stream/socket/fs string write) gets 4–5× faster. A 64 KiB JSON-RPC round trip over a child's stdio (stringify → write → readline → parse, non-ASCII payload) drops from 1.28 ms to 0.90 ms with only the parent patched, 0.67 ms with both ends.

string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=1024      ***  1125.82 %  ±8.25%
string_decoder/string-decoder.js encoding='utf8' inLen=128  chunkLen=1024      ***   322.79 %  ±2.28%
string_decoder/string-decoder.js encoding='utf8' inLen=32   chunkLen=1024      ***    98.53 %  ±1.37%
string_decoder/string-decoder.js encoding='utf8' inLen=1024 chunkLen=16        ***     5.15 %  ±0.58%
buffers/buffer-write-string-utf8.js (new)  len=65536 chars='two-byte'                 ***   391.17 %
buffers/buffer-write-string-utf8.js        len=65536 chars='two-byte-lone-surrogate'  ***   314.76 %
buffers/buffer-write-string-utf8.js        len=256   chars='two-byte'                 ***   289.28 %
buffers/buffer-write-string-utf8.js        len=256   chars='two-byte-astral'          ***   228.97 %
one-byte strings / other encodings                                                             ~0 %  n.s.

(Linux x64, benchmark/compare.js, 30 runs, significance as in compare.R.)

Two commits:

string_decoder: decode UTF-8 via StringBytes::Encode - the decoder built strings with String::NewFromUtf8(); buffer.toString() already goes through StringBytes::Encode(), which uses simdutf. MakeString() now calls the same function (the kMaxLengthERR_STRING_TOO_LONG check stays in front), so the decoder returns exactly what toString() returns for the same bytes. The partial-character bookkeeping is untouched.

src: use simdutf for two-byte strings in UTF-8 writes - StringBytes::Write(UTF8) used String::WriteUtf8V2() for two-byte strings. For strings longer than 32 code units it now validates with simdutf::validate_utf16() (lone surrogates are repaired into a scratch buffer with to_well_formed_utf16(), so the output stays byte-identical to V8's kReplaceInvalidUtf8) and transcodes with convert_utf16_to_utf8() - but only when the destination is known to fit (buflen >= 3 × length or >= utf8_length_from_utf16()); otherwise it falls back to V8, so partial writes truncate at the same character boundary as before. Shorter strings and one-byte strings keep the V8 path.

Tests: test-string-decoder-utf8-large.js (new: large inputs, chunk boundaries inside multi-byte sequences, invalid sequences vs toString()); test-buffer-write-utf8-two-byte.js (new: independent reference encoder; BMP/astral/lone surrogates at start/middle/end; exact-fit, one-short and 3× destinations; partial writes). Existing string_decoder, buffer, stream, net, http, child_process and readline suites pass.


Disclosure: the code, test, measurements and this description were written by Claude Code, directed and reviewed by @codebytere.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/performance

@nodejs-github-bot nodejs-github-bot added buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. needs-ci PRs that need a full CI run. labels Aug 16, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
@codebytere
codebytere force-pushed the perf/src-simdutf-utf8-transcoding branch from ec59ddb to 0c89308 Compare August 16, 2026 15:51
@codebytere codebytere added request-ci Add this label to start a Jenkins CI on a PR. and removed needs-ci PRs that need a full CI run. labels Aug 16, 2026
@codecov

codecov Bot commented Aug 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.11%. Comparing base (30bff4a) to head (0c89308).
⚠️ Report is 38 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #65324      +/-   ##
==========================================
- Coverage   90.13%   90.11%   -0.03%     
==========================================
  Files         752      752              
  Lines      251568   251578      +10     
  Branches    47270    47273       +3     
==========================================
- Hits       226759   226706      -53     
- Misses      16168    16216      +48     
- Partials     8641     8656      +15     
Files with missing lines Coverage Δ
src/string_bytes.cc 75.23% <100.00%> (+1.49%) ⬆️
src/string_decoder.cc 92.17% <100.00%> (-0.34%) ⬇️

... and 37 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@codebytere
codebytere requested review from addaleax and anonrig August 16, 2026 18:37
@github-actions github-actions Bot removed the request-ci Add this label to start a Jenkins CI on a PR. label Aug 16, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Comment thread src/string_bytes.cc
// 3 * length is what StorageSize() hands most callers; only compute
// the exact length when the buffer is smaller than that.
if (buflen >= 3 * length ||
buflen >= simdutf::utf8_length_from_utf16(data, length)) {

@lemire lemire Aug 16, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have simdutf::utf8_length_from_utf16_with_replacement that would be appropriate in this function.

Note that we added convert_utf16_to_utf8_with_replacement in arelease this year. (But this code is still quite fine.)

@lemire

lemire commented Aug 16, 2026

Copy link
Copy Markdown
Member

@codebytere I might send you a PR later today. Hold on a bit.

@lemire

lemire commented Aug 16, 2026

Copy link
Copy Markdown
Member

@codebytere The simdutf version might not be ready yet. So I still recommend merging this PR. It is good work. We might be able to improve it later, but this should not stop this PR.

@github-actions

github-actions Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

@codebytere codebytere added the commit-queue Add this label to land a pull request using GitHub Actions. label Aug 17, 2026
@codebytere

Copy link
Copy Markdown
Member Author

@jasnell is commit queue broken?

@jasnell

jasnell commented Aug 18, 2026

Copy link
Copy Markdown
Member

Entirely possible

@nodejs-github-bot nodejs-github-bot added commit-queue-failed An error occurred while landing this pull request using GitHub Actions. and removed commit-queue Add this label to land a pull request using GitHub Actions. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator
Commit Queue failed
- Loading data for nodejs/node/pull/65324
✔  Done loading data for nodejs/node/pull/65324
----------------------------------- PR info ------------------------------------
Title      src: use simdutf for UTF-8 ⇄ UTF-16 transcoding in StringDecoder and StringBytes::Write (#65324)
Author     Shelley Vohr <shelley.vohr@gmail.com> (@codebytere)
Branch     codebytere:perf/src-simdutf-utf8-transcoding -> nodejs:main
Labels     buffer, c++, commit-queue
Commits    2
 - string_decoder: decode UTF-8 via StringBytes::Encode
 - src: use simdutf for two-byte strings in UTF-8 writes
Committers 1
 - Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
------------------------------ Generated metadata ------------------------------
PR-URL: https://github.com/nodejs/node/pull/65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
--------------------------------------------------------------------------------
   ℹ  This PR was created on Sun, 16 Aug 2026 15:37:44 GMT
   ✔  Approvals: 4
   ✔  - Yagiz Nizipli (@anonrig) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947047263
   ✔  - Gürgün Dayıoğlu (@gurgunday): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947063127
   ✔  - Daniel Lemire (@lemire): https://github.com/nodejs/node/pull/65324#pullrequestreview-4947151736
   ✔  - James M Snell (@jasnell) (TSC): https://github.com/nodejs/node/pull/65324#pullrequestreview-4948248917
   ✔  Last GitHub CI successful
   ℹ  Last Full PR CI on 2026-08-16T18:58:59Z: https://ci.nodejs.org/job/node-test-pull-request/75896/
- Querying data for job/node-test-pull-request/75896/
✔  Build data downloaded
   ✔  Last Jenkins CI successful
--------------------------------------------------------------------------------
   ✔  No git cherry-pick in progress
   ✔  No git am in progress
   ✔  No git rebase in progress
--------------------------------------------------------------------------------
- Bringing origin/main up to date...
From https://github.com/nodejs/node
 * branch                  main       -> FETCH_HEAD
✔  origin/main is now up-to-date
- Downloading patch for 65324
From https://github.com/nodejs/node
 * branch                  refs/pull/65324/merge -> FETCH_HEAD
✔  Fetched commits as 0d84f29fa77a..0c89308bba0c
--------------------------------------------------------------------------------
[main 308a46181c] string_decoder: decode UTF-8 via StringBytes::Encode
 Author: Shelley Vohr <shelley.vohr@gmail.com>
 Date: Sat Aug 15 22:09:21 2026 +0000
 2 files changed, 94 insertions(+), 15 deletions(-)
 create mode 100644 test/parallel/test-string-decoder-utf8-large.js
[main 619cb0dd4c] src: use simdutf for two-byte strings in UTF-8 writes
 Author: Shelley Vohr <shelley.vohr@gmail.com>
 Date: Sat Aug 15 22:34:28 2026 +0000
 3 files changed, 193 insertions(+), 1 deletion(-)
 create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
 create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
   ✔  Patches applied
There are 2 commits in the PR. Attempting autorebase.
(node:381) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true can lead to security vulnerabilities, as the arguments are not escaped, only concatenated.
(Use `node --trace-deprecation ...` to show where the warning was created)
Rebasing (2/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
string_decoder: decode UTF-8 via StringBytes::Encode

StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 1055900345] string_decoder: decode UTF-8 via StringBytes::Encode
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:09:21 2026 +0000
2 files changed, 94 insertions(+), 15 deletions(-)
create mode 100644 test/parallel/test-string-decoder-utf8-large.js
Rebasing (3/4)
Rebasing (4/4)
Executing: git node land --amend --yes
--------------------------------- New Message ----------------------------------
src: use simdutf for two-byte strings in UTF-8 writes

StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>

[detached HEAD 55c27c9b17] src: use simdutf for two-byte strings in UTF-8 writes
Author: Shelley Vohr <shelley.vohr@gmail.com>
Date: Sat Aug 15 22:34:28 2026 +0000
3 files changed, 193 insertions(+), 1 deletion(-)
create mode 100644 benchmark/buffers/buffer-write-string-utf8.js
create mode 100644 test/parallel/test-buffer-write-utf8-two-byte.js
Successfully rebased and updated refs/heads/main.

ℹ Add commit-queue-squash label to land the PR as one commit, or commit-queue-rebase to land as separate commits.

https://github.com/nodejs/node/actions/runs/32156233387

@aduh95 aduh95 added commit-queue Add this label to land a pull request using GitHub Actions. commit-queue-rebase Add this label to allow the Commit Queue to land a PR in several commits. and removed commit-queue-failed An error occurred while landing this pull request using GitHub Actions. labels Aug 18, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Landed in 29a083c...55e4ca3

nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringDecoder used v8::String::NewFromUtf8() for UTF-8, while
Buffer#toString() goes through StringBytes::Encode(), which has
simdutf-backed ASCII, Latin-1 and UTF-16 paths and only falls back
to NewFromUtf8() for input that contains invalid sequences. Route
the decoder through the same function, so streams with
setEncoding('utf8') and readline decode at the same speed as
Buffer#toString(). U+FFFD replacement is unchanged because invalid
input still ends up in NewFromUtf8(), and the ERR_STRING_TOO_LONG
check is kept explicit so over-long input fails as before.

benchmark/string_decoder/string-decoder.js (encoding=utf8) and a
readline-over-pipe workload improve by 2-3x for chunks >= 1 KiB;
64 KiB newline-delimited JSON round trips over child stdio improve
by ~30% on the reading side alone.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
nodejs-github-bot pushed a commit that referenced this pull request Aug 18, 2026
StringBytes::Write() already used simdutf to encode one-byte strings
as UTF-8 but sent every two-byte (UTF-16) string through
v8::String::WriteUtf8V2(), which is several times slower. That path
is behind Buffer.from(string), buf.write(), fs.write*() with string
data and every string written to a libuv stream, and JSON.stringify()
output is a two-byte string as soon as any value in the payload is
outside Latin-1.

Encode two-byte strings with simdutf as well whenever their UTF-8 form
is guaranteed to fit in the target: well-formed input is converted
directly, and input with unpaired surrogates is converted from a copy
passed through simdutf::to_well_formed_utf16(), which replaces each
unpaired surrogate with U+FFFD exactly like kReplaceInvalidUtf8 (this
mirrors what TextEncoder already does). Writes that have to truncate
at a character boundary keep using WriteUtf8V2(), so their output is
byte-for-byte unchanged, and so do strings of up to 32 code units, for
which V8 is already as fast (the same threshold TextEncoder uses).

buf.write() of a 2 KiB two-byte string improves ~5x (astral-heavy and
lone-surrogate strings ~3.5x and ~5x), Buffer.from() of a 64 KiB JSON
string ~2.7x; one-byte strings are unaffected.

Signed-off-by: Shelley Vohr <shelley.vohr@gmail.com>
PR-URL: #65324
Reviewed-By: Yagiz Nizipli <yagiz@nizipli.com>
Reviewed-By: Gürgün Dayıoğlu <hey@gurgun.day>
Reviewed-By: Daniel Lemire <daniel@lemire.me>
Reviewed-By: James M Snell <jasnell@gmail.com>
@nodejs-github-bot nodejs-github-bot removed the commit-queue Add this label to land a pull request using GitHub Actions. label Aug 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

buffer Issues and PRs related to the buffer subsystem. c++ Issues and PRs that require attention from people who are familiar with C++. commit-queue-rebase Add this label to allow the Commit Queue to land a PR in several commits.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants