Skip to content

test: add matched native Compose overhead benchmark - #180

Merged
roodboi merged 3 commits into
nextfrom
test/native-compose-overhead
Oct 8, 2026
Merged

roodboi merged 3 commits into
nextfrom
test/native-compose-overhead

Conversation

@roodboi

@roodboi roodboi commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

This adds a preview-first benchmark for native authored configuration versus legacy Hack configuration and Compose, using the same frozen CLI/compiler, engine, cached image and app/initializer/storage policy. It measures warm up, ps, exec, restart and down in eight alternating paired rounds for one- and two-project cohorts. Common readiness, data and route checks, process-tree CPU accounting, declared regression/noise gates and exact owned cleanup remain part of the comparison. Native compilation and ownership proof stay in the timed commands.

An optional private receipt selects a separately qualified proxy with no published ports, two ephemeral tmpfs and only a readonly engine socket bind. Both lanes use the same bounded in-proxy HTTP oracle and exact proxy/network/CA binding checks. The ordinary host-port-80 mode remains available. Raw results stay private. The fixture omits custom networks, host hooks, material files and builds; this compares image-only authoring overhead through the current CLI, not installed v4 or application/VM performance.

The module also provides a bounded direct-stdin sampler for compiler protocol, compile and metadata-plan costs. It pins the executable, input and independently qualified canonical output, uses anonymous regular stdin and one-ms wait4 polling, and preserves failed observations with their actual metrics. Real-child controls cover cancellation, reaping, output bounds, tamper and successful exits observed beyond the unchanged deadline. Importing the module collects no samples or engine effects.

Validation: head 52e66e61 is normally reconciled onto next dccd3f97; all four benchmark-owned source/test/doc/CI files retain the bytes qualified at 7085c4e7. All 65 offline controls, CLI typecheck, Python syntax, workflow validation and diff checks pass. Independent cumulative source review cleared this exact head; all 12 exact-head CI jobs pass. Automated Graphite review is pending. Historical full local CLI proof (3194 passed, 73 skipped, no failures) is preserved; the unchanged whole suite was not repeated locally. Actual measurements below use the immutable 7085c4e7 CLI (Bun 1.4.2) and matching Rust 1.97.1 optimized release compiler, not a newly measured reconciled binary.

Actual M3 qualification completed all 160 declared observations and six fixture cleanups. Exact proxy retirement and both owner and independent full original-inventory postflights passed. The comparison exited 1 and remains unqualified: 129 observations and 72 paired rounds were flagged for host load, leaving fewer than eight usable pairs per action. Every row and the original pre-effect admission refusal are retained; no samples were replaced and no thresholds changed. This is no evidence of a performance improvement.

The separate quiet compiler cohort passed all 30 observations (six warmups, 24 measured) and its declared absolute budgets. Median direct compile cost was 5.25 ms elapsed and 2.53 ms CPU; metadata-plan cost was 5.40 ms elapsed and 2.69 ms CPU. These standalone costs do not estimate integrated CLI or VM overhead.

A separately reviewed single M3 lifecycle attribution run completed up→restart→down, verified readiness/new boot/data, exact resource retirement and both full original-inventory postflights. Actual Docker launch counts were 344/365/91; Docker child CPU was 12.88/13.52/3.16 seconds. Startup and restart each invoked the compiler 170 times, including 136 protocol handshakes, using about 0.42 seconds of compiler CPU per action. The largest observed category was docker info: 60 calls per startup/restart used about 4.84/4.88 seconds of Docker child CPU. These counts identify repeated engine-process queries as an optimization target while preserving ownership and freshness checks.

This trace used a private category-only shim that recorded no arguments, environment or I/O. Its interference was substantial: instrumented CLI-tree CPU was 71.10/74.55/13.65 seconds, while recorded shim CPU lower bounds were 54.14/56.66/9.61 seconds. Consequently the trace supplies subprocess counts and child CPU attribution only; it is not an ordinary CLI performance comparison, replacement sample or speedup claim.

The original 7085c4e7 CI failure is preserved: the macOS Intel compiler job exceeded inherited five-second outer file-test watchdogs. Current next includes the narrowly reviewed outer-test-budget corrections, including uncertain after-hook coverage; product deadlines and assertions remain unchanged. All 12 reconciled-head CI jobs pass, and the PR is ready for review. Automated Graphite review remains a separate merge gate. No release publication is included.

@roodboi
roodboi force-pushed the test/native-compose-overhead branch 8 times, most recently from f7db65e to 7085c4e Compare October 8, 2026 10:28
@roodboi
roodboi force-pushed the test/native-compose-overhead branch from 7085c4e to 52e66e6 Compare October 8, 2026 12:39
@roodboi
roodboi marked this pull request as ready for review October 8, 2026 13:01
@roodboi
roodboi merged commit 521a237 into next Oct 8, 2026
13 checks passed
@roodboi
roodboi deleted the test/native-compose-overhead branch October 8, 2026 13:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant