Repository navigation
test: add matched native Compose overhead benchmark - #180
Merged
Merged
Conversation
roodboi
force-pushed
the
test/native-compose-overhead
branch
8 times, most recently
from
October 8, 2026 10:28
f7db65e to
7085c4e
Compare
roodboi
force-pushed
the
test/native-compose-overhead
branch
from
October 8, 2026 12:39
7085c4e to
52e66e6
Compare
roodboi
marked this pull request as ready for review
October 8, 2026 13:01
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds a preview-first benchmark for native authored configuration versus legacy Hack configuration and Compose, using the same frozen CLI/compiler, engine, cached image and app/initializer/storage policy. It measures warm up, ps, exec, restart and down in eight alternating paired rounds for one- and two-project cohorts. Common readiness, data and route checks, process-tree CPU accounting, declared regression/noise gates and exact owned cleanup remain part of the comparison. Native compilation and ownership proof stay in the timed commands.
An optional private receipt selects a separately qualified proxy with no published ports, two ephemeral tmpfs and only a readonly engine socket bind. Both lanes use the same bounded in-proxy HTTP oracle and exact proxy/network/CA binding checks. The ordinary host-port-80 mode remains available. Raw results stay private. The fixture omits custom networks, host hooks, material files and builds; this compares image-only authoring overhead through the current CLI, not installed v4 or application/VM performance.
The module also provides a bounded direct-stdin sampler for compiler protocol, compile and metadata-plan costs. It pins the executable, input and independently qualified canonical output, uses anonymous regular stdin and one-ms wait4 polling, and preserves failed observations with their actual metrics. Real-child controls cover cancellation, reaping, output bounds, tamper and successful exits observed beyond the unchanged deadline. Importing the module collects no samples or engine effects.
Validation: head
52e66e61is normally reconciled onto nextdccd3f97; all four benchmark-owned source/test/doc/CI files retain the bytes qualified at7085c4e7. All 65 offline controls, CLI typecheck, Python syntax, workflow validation and diff checks pass. Independent cumulative source review cleared this exact head; all 12 exact-head CI jobs pass. Automated Graphite review is pending. Historical full local CLI proof (3194 passed, 73 skipped, no failures) is preserved; the unchanged whole suite was not repeated locally. Actual measurements below use the immutable7085c4e7CLI (Bun 1.4.2) and matching Rust 1.97.1 optimized release compiler, not a newly measured reconciled binary.Actual M3 qualification completed all 160 declared observations and six fixture cleanups. Exact proxy retirement and both owner and independent full original-inventory postflights passed. The comparison exited 1 and remains unqualified: 129 observations and 72 paired rounds were flagged for host load, leaving fewer than eight usable pairs per action. Every row and the original pre-effect admission refusal are retained; no samples were replaced and no thresholds changed. This is no evidence of a performance improvement.
The separate quiet compiler cohort passed all 30 observations (six warmups, 24 measured) and its declared absolute budgets. Median direct compile cost was 5.25 ms elapsed and 2.53 ms CPU; metadata-plan cost was 5.40 ms elapsed and 2.69 ms CPU. These standalone costs do not estimate integrated CLI or VM overhead.
A separately reviewed single M3 lifecycle attribution run completed up→restart→down, verified readiness/new boot/data, exact resource retirement and both full original-inventory postflights. Actual Docker launch counts were 344/365/91; Docker child CPU was 12.88/13.52/3.16 seconds. Startup and restart each invoked the compiler 170 times, including 136 protocol handshakes, using about 0.42 seconds of compiler CPU per action. The largest observed category was docker info: 60 calls per startup/restart used about 4.84/4.88 seconds of Docker child CPU. These counts identify repeated engine-process queries as an optimization target while preserving ownership and freshness checks.
This trace used a private category-only shim that recorded no arguments, environment or I/O. Its interference was substantial: instrumented CLI-tree CPU was 71.10/74.55/13.65 seconds, while recorded shim CPU lower bounds were 54.14/56.66/9.61 seconds. Consequently the trace supplies subprocess counts and child CPU attribution only; it is not an ordinary CLI performance comparison, replacement sample or speedup claim.
The original
7085c4e7CI failure is preserved: the macOS Intel compiler job exceeded inherited five-second outer file-test watchdogs. Current next includes the narrowly reviewed outer-test-budget corrections, including uncertain after-hook coverage; product deadlines and assertions remain unchanged. All 12 reconciled-head CI jobs pass, and the PR is ready for review. Automated Graphite review remains a separate merge gate. No release publication is included.