Reduce Ruby overhead in get_multi reply parsing and key routing - #1169
Merged
Merged
Conversation
radixdev
marked this pull request as ready for review
September 29, 2026 16:05
- Parse pipelined VA reply headers in place instead of splitting them into tokens (multi-server get_multi) - Parse single-server get_multi VA lines in place as well - Find a key's continuum entry through a bucket table instead of a binary search, and try the key's own server before the failover loop Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
petergoldstein
force-pushed
the
perf/get-multi-100
branch
from
October 2, 2026 03:04
b343180 to
f3dfb2d
Compare
Owner
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR lowers the Ruby cost of
get_multiin two places, parsing replies and routing keys to servers. A 100-keyget_multiover loopback gets 7% faster than Dalli 2.7.11 on 4 servers and 6% faster on 16. Onmaintoday it's 16% and 9% slower. Results and wire bytes don't change.Why
On a 4-server ring, a 100-key
get_multitook about 287µs onmain, against about 250µs on 2.7.11 (Braze's 2.7.11 fork). CPU time and wall time were about the same, so the gap is Ruby work, not waiting.getkq33.9µs, metav f k q s34.4µs, and withouts34.3µs.getk_response_from_buffer. EachVAheader is copied into a string and split into about 5 token strings, thendelete_prefix!andto_irun on each flag. That's about 12 allocations per key.Ring#server_for_hash_keyruns a binary search of the continuum, which is about 10 Ruby block calls per key (12 on 16 servers). 2.7.11 does the same, so this wasn't part of the gap, but it's a cheap win.Meta#parse_multi_get_value) splits eachVAline into tokens in the same way.Changes
ResponseProcessor#va_response_from_bufferreads aVAheader in place, without a header string, token array or flag strings.getk_response_from_bufferuses it forVAlines.bflag.sflag or the size is zero, it returnsniland the existing token path handles the line, with the same behavior as before.Ring#server_for_hash_keylooks the hash up in a table with about one bucket per continuum entry. Each bucket holds the index of the first entry at or after its start, so the lookup steps over a few entries instead of running a binary search.continuum=: 1,024 buckets for 4 servers, 4,096 for 16.server_from_continuumtries the key's own server before the failover loop. The failover order is the same as before.server_alive?looks up its per-call cache without a block.Meta#parse_multi_get_value(the single-server path) reads the size, flags and key straight from the line withsize_from_va_line,bitflags_from_va_lineand the newkey_from_va_line.Benchmarks
Ruby 3.4.10 and memcached 1.6.38 in Docker on Apple Silicon. memcached and Ruby are pinned to separate CPUs, and each number is the fastest slice over 3 rounds run in rotating order. Values are 100-byte strings. The comparison is
main, this PR, and Braze's 2.7.11 fork, as a baseline.Loopback, µs per
get_multi:mainOn 4 servers with 100 keys, allocations per call go from 1,500 to 1,104 (2.7.11 makes 1,318).
About 300µs round trip (
tc netem, 100µs per packet):mainWith latency, this PR's CPU time per call matches 2.7.11's (257µs vs 260µs). The rest of the difference is time spent waiting, about 10µs more per call than 2.7.11.
mainwaits the same amount without this PR, so it comes from how the requests and replies are timed, not from Ruby work. I haven't tracked it down.In a parser microbenchmark on 25 replies, the in-place path takes about 29µs against about 40µs for the token path, and makes half the allocations. Both figures include
Marshal.load.Testing
bundle exec rake: 849 runs, 0 failures, 0 errors.bbefore and afterk, a missingfork, junk tokens and junk after digits, partial headers and bodies, a zero size, and a missings. There's also a check that a normal hit takes the in-place path.key_from_va_line, including base64 keys and a key namedb.bundle exec rubocop: no offenses.This PR was generated with the assistance of Claude Code (Anthropic). The benchmark and test results above come from real runs.
🤖 Generated with Claude Code