Skip to content

Reduce Ruby overhead in get_multi reply parsing and key routing - #1169

Merged
petergoldstein merged 2 commits into
petergoldstein:mainfrom
radixdev:perf/get-multi-100
Oct 2, 2026
Merged

petergoldstein merged 2 commits into
petergoldstein:mainfrom
radixdev:perf/get-multi-100

Conversation

@radixdev

Copy link
Copy Markdown
Contributor

Summary

This PR lowers the Ruby cost of get_multi in two places, parsing replies and routing keys to servers. A 100-key get_multi over loopback gets 7% faster than Dalli 2.7.11 on 4 servers and 6% faster on 16. On main today it's 16% and 9% slower. Results and wire bytes don't change.

Why

On a 4-server ring, a 100-key get_multi took about 287µs on main, against about 250µs on 2.7.11 (Braze's 2.7.11 fork). CPU time and wall time were about the same, so the gap is Ruby work, not waiting.

  • memcached isn't the cause. With raw sockets and no Dalli code, memcached 1.6.38 answers a batch in the same time in every format. For 25 keys: binary getkq 33.9µs, meta v f k q s 34.4µs, and without s 34.3µs.
  • Reply parsing. The profile shows most of the gap in getk_response_from_buffer. Each VA header is copied into a string and split into about 5 token strings, then delete_prefix! and to_i run on each flag. That's about 12 allocations per key.
  • Routing. Ring#server_for_hash_key runs a binary search of the continuum, which is about 10 Ruby block calls per key (12 on 16 servers). 2.7.11 does the same, so this wasn't part of the gap, but it's a cheap win.
  • The single-server path (Meta#parse_multi_get_value) splits each VA line into tokens in the same way.

Changes

  • ResponseProcessor#va_response_from_buffer reads a VA header in place, without a header string, token array or flag strings.
    • getk_response_from_buffer uses it for VA lines.
    • It returns the same values as the token path, including "the first of a repeated flag wins" and the b flag.
    • When there's no s flag or the size is zero, it returns nil and the existing token path handles the line, with the same behavior as before.
  • Ring#server_for_hash_key looks the hash up in a table with about one bucket per continuum entry. Each bucket holds the index of the first entry at or after its start, so the lookup steps over a few entries instead of running a binary search.
    • The table is built in continuum=: 1,024 buckets for 4 servers, 4,096 for 16.
    • server_from_continuum tries the key's own server before the failover loop. The failover order is the same as before.
    • server_alive? looks up its per-call cache without a block.
  • Meta#parse_multi_get_value (the single-server path) reads the size, flags and key straight from the line with size_from_va_line, bitflags_from_va_line and the new key_from_va_line.

Benchmarks

Ruby 3.4.10 and memcached 1.6.38 in Docker on Apple Silicon. memcached and Ruby are pinned to separate CPUs, and each number is the fastest slice over 3 rounds run in rotating order. Values are 100-byte strings. The comparison is main, this PR, and Braze's 2.7.11 fork, as a baseline.

Loopback, µs per get_multi:

servers keys 2.7.11 main this PR
1 3 44.8 43.2 40.4
1 100 221.6 226.8 213.9
4 3 45.1 43.4 44.8
4 20 87.0 95.7 84.1
4 100 260.6 301.5 243.3
16 20 149.5 163.3 148.4
16 100 363.7 397.4 342.3

On 4 servers with 100 keys, allocations per call go from 1,500 to 1,104 (2.7.11 makes 1,318).

About 300µs round trip (tc netem, 100µs per packet):

servers keys 2.7.11 main this PR
4 20 362.3 379.1 364.3
4 100 517.2 580.8 526.0
16 20 389.2 391.2 379.4
16 100 530.5 578.3 533.1

With latency, this PR's CPU time per call matches 2.7.11's (257µs vs 260µs). The rest of the difference is time spent waiting, about 10µs more per call than 2.7.11. main waits the same amount without this PR, so it comes from how the requests and replies are timed, not from Ruby work. I haven't tracked it down.

In a parser microbenchmark on 25 replies, the in-place path takes about 29µs against about 40µs for the token path, and makes half the allocations. Both figures include Marshal.load.

Testing

  • bundle exec rake: 849 runs, 0 failures, 0 errors.
  • New tests:
    • The in-place parser against the token path on 15 header shapes: flags in any order, repeated flags, b before and after k, a missing f or k, junk tokens and junk after digits, partial headers and bodies, a zero size, and a missing s. There's also a check that a normal hit takes the in-place path.
    • key_from_va_line, including base64 keys and a key named b.
    • The bucket lookup against a binary search of the continuum on rings of 2, 16 and 50 weighted servers, probed at every entry's value, value−1 and value+1, at 0 and 2³²−1, and at random hashes.
    • Failover with a dead server picks the same server as the previous order.
  • Separately, I compared both lookups on about 1M hashes across 2 to 50 servers: 0 mismatches.
  • bundle exec rubocop: no offenses.

This PR was generated with the assistance of Claude Code (Anthropic). The benchmark and test results above come from real runs.

🤖 Generated with Claude Code

@radixdev
radixdev marked this pull request as ready for review September 29, 2026 16:05
radixdev and others added 2 commits October 1, 2026 22:54
- Parse pipelined VA reply headers in place instead of splitting them
  into tokens (multi-server get_multi)
- Parse single-server get_multi VA lines in place as well
- Find a key's continuum entry through a bucket table instead of a
  binary search, and try the key's own server before the failover loop

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@petergoldstein
petergoldstein merged commit 57e6010 into petergoldstein:main Oct 2, 2026
17 checks passed
@petergoldstein

Copy link
Copy Markdown
Owner

Thanks, merged. I rebased it onto main after #1170. The only conflict was in the tests, and both sets are kept. Your in-place VA parser hands zero-size replies back to the token path, so it picks up #1170's empty-value fix. This ships in 5.2.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants