diff --git a/docs/research/CACHE_USAGE_CONTRACT_2026-09-11.md b/docs/research/CACHE_USAGE_CONTRACT_2026-09-11.md new file mode 100644 index 00000000..d23e3e73 --- /dev/null +++ b/docs/research/CACHE_USAGE_CONTRACT_2026-09-11.md @@ -0,0 +1,76 @@ +# Cache usage fields and synthetic replay boundary + +- Status: validated for the named source inspection and deterministic synthetic tests only; no provider experiment or comparative Session replay. +- Created and verified: 2026-09-11. +- Source boundary: OpenPI `5bf2fe29e52801d79826c2eb573be403f53285e5` and the scoped tests accompanying this record. Pi `@earendil-works/pi-ai` 0.85.1 is locked in `bun.lock`; its [published metadata](https://registry.npmjs.org/@earendil-works/pi-ai/0.85.1) and `v0.85.1` tag identify source commit `d981de1229ef899957bbe968bc8dcda02a21f477`. The installed package's relevant compiled mappings were also inspected. +- Related: [Issue #156](https://github.com/openpi-dev/openpi/issues/156), [diagnostic core PR #372](https://github.com/openpi-dev/openpi/pull/372), and the [remaining-evidence boundary](https://github.com/openpi-dev/openpi/issues/156#issuecomment-5594640168). This record does not close #156. +- Supersedes: none. This is a new source snapshot, not a reinterpretation of the older Pi revision cited in the Issue. + +## What OpenPI can observe + +Pi owns provider normalization. OpenPI's existing tracker consumes normalized `Usage` from assistant `turn_end` events other than errors and aborts, when a turn identity is available; it does not inspect raw API usage or determine provider billing. The [Pi type](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/types.ts#L383) carries required numeric counters but no per-counter presence flags, cache key, expiry timestamp, or invalidation reason. + +| Normalized field | Use and limitation at this source boundary | +| --- | --- | +| `input` | The adapter's uncached input component, with the provider-specific mappings below. Do not add raw `prompt_tokens` or `input_tokens` again. | +| `cacheRead` | Reported cache-read tokens after adapter normalization. A positive value supports a read observation; zero alone cannot distinguish an omitted field, unsupported reporting, or an actual zero. | +| `cacheWrite` | Reported cache-creation/write tokens where mapped. Zero is not proof that no provider-side caching happened. | +| `cacheWrite1h` | Optional subset of `cacheWrite`, populated by the inspected Anthropic path. It is not an additional prompt component and does not identify why a later read disappeared. | +| `output`, `reasoning` | Output accounting; optional `reasoning` is already a subset of `output`. Neither belongs in the prompt-cache denominator. | +| `totalTokens` | Adapter-reported or computed total. OpenPI derives prompt tokens from `input + cacheRead + cacheWrite`, not from this field. | +| `cost.*` | Pi cost accounting, not a cache-causality signal. The inspected adapters call [Pi's model-rate calculation](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/models.ts#L891); these numbers are not an independently verified invoice or proof of savings. | + +## Provider/API field matrix + +These are source-verified adapter assignments, not assertions that every endpoint emits every field. Optional raw counters generally default to zero. Adapter/API identity and OpenPI's provider-name classification are separate: sharing an API does not promote a provider to `explicit-prefix`. + +| Pi adapter and source | `input` mapping | `cacheRead` mapping | `cacheWrite` mapping and unknown boundary | +| --- | --- | --- | --- | +| [Anthropic Messages](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/anthropic-messages.ts#L606) | `input_tokens` | `cache_read_input_tokens` | `cache_creation_input_tokens`; `cache_creation.ephemeral_1h_input_tokens` initializes the optional write subset. Later message deltas update the reported input/read/write counters when present. No expiry cause is supplied to OpenPI. | +| [OpenAI-compatible Completions](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/openai-completions.ts#L1507) | `max(0, prompt_tokens - read - write)` | First non-nullish value of `prompt_tokens_details.cached_tokens`, `prompt_cache_hit_tokens`, `cached_tokens`; otherwise zero | `prompt_tokens_details.cache_write_tokens`, default zero. Mapping a compatibility field does not establish which upstream provider performed a write. | +| [OpenAI Responses shared handler](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/openai-responses-shared.ts#L559) | `max(0, input_tokens - read - write)` | `input_tokens_details.cached_tokens` | `input_tokens_details.cache_write_tokens`, default zero. [Azure](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/azure-openai-responses.ts#L131) and [Codex](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/openai-codex-responses.ts#L660) also use this handler. | +| [Google Generative AI](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/google-generative-ai.ts#L223) | `promptTokenCount - cachedContentTokenCount` from `usageMetadata`, with absent values zero | `cachedContentTokenCount` | Always zero in this mapping. A zero write count cannot establish that caching is disabled or that creation was free. | +| [Google Vertex](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/google-vertex.ts#L240) | Same Google metadata subtraction | `cachedContentTokenCount` | Always zero in this mapping. Despite the similar adapter fields, `google-vertex` remains an unknown provider in OpenPI's current classifier. | +| [Bedrock Converse](https://github.com/earendil-works/pi/blob/d981de1229ef899957bbe968bc8dcda02a21f477/packages/ai/src/api/bedrock-converse-stream.ts#L685) | `inputTokens` | `cacheReadInputTokens` | `cacheWriteInputTokens`. These are direct assignments; `amazon-bedrock` remains unknown to OpenPI's classifier even when its underlying model is Anthropic. | + +Other adapters, custom endpoints, and account-specific billing have not been verified by this matrix. In particular, the current OpenPI name allowlist retains `google-antigravity`, `google-gemini-cli`, and `openai-responses`; their presence in that list is not evidence of a corresponding built-in provider at the pinned Pi revision. + +## Existing detector contract, not a new policy + +The [OpenPI classifier](../../extensions/model-info/cache-diagnostics.ts) has these provider-name boundaries: + +- Only normalized `anthropic` is `explicit-prefix`. +- `azure-openai-responses`, `google`, `google-antigravity`, `google-gemini-cli`, `openai`, `openai-codex`, and `openai-responses` are `implicit-best-effort`. +- Every other name is `unknown`, including compatible proxies. This record does not expand the allowlist. + +The first accepted observation is `first-turn`, even if it contains a positive read. A subsequent zero-read turn is `cold` when the previous read was below 2,048 tokens. At or above that boundary, a zero read becomes `miss-after-warm-prefix` only for `explicit-prefix`; otherwise it remains `unknown`. A positive read lower than the previous read is `partial-hit`; other positive reads are `warm`. + +For `miss-after-warm-prefix`, `reprocessedTokens` is the current normalized `input`, not total prompt tokens, cache writes, or a proven count of an identical prefix processed twice. The current detector does not require `cacheWrite > 0`, unlike the OMP predicate linked in #156. Model, thinking, tool, system-prompt, compaction, and branch changes are correlations only. Every result has `evidence: "observation"` and `verifiedCause: null`. + +The [extension event seam](../../extensions/model-info/index.ts) skips error/aborted replies without replacing the last accepted baseline. Compaction and branch movement add pending correlations; they do not reset that baseline. Session start and shutdown reset it. A new Session start does not reconstruct diagnostic continuity from saved history, even though the separate [active-branch metric](../../extensions/model-info/session-metrics.ts) reads persisted usage. That metric includes assistant, tool-result, compaction, and branch-summary usage; it is not the same population as diagnostic assistant turns, a whole Session tree, or all child/account costs. + +The tracker normalizes non-finite or nonpositive prompt counters to zero. This is defensive arithmetic, not recovered evidence that the provider actually reported zero. Its observation input has no timestamp or retention metadata. A long delay, a read of zero, or a reported one-hour write subset therefore cannot establish `TTL expiry` as a separate verified kind. + +## Reproducible synthetic evidence + +The fixtures are small, hand-authored TypeScript vectors in the existing tests. They contain no user transcript, API response capture, credentials, or model invocation. They exercise the shipped tracker and extension handlers, not a reimplementation of the detector. + +| Test surface | Additional assurance | +| --- | --- | +| [`cache-diagnostics.test.ts`](../../tests/extensions/model-info/cache-diagnostics.test.ts) | Six observations for each of the seven implicit names and four unknown examples retain unknown/cold/partial-hit boundaries, even with positive write counters. Exact 2,047/2,048 threshold, prompt-component arithmetic, excluded subsets, immutable input, and malformed-number handling are checked separately. | +| [`index.test.ts`](../../tests/extensions/model-info/index.test.ts) | Seven-turn event replay combines warm/cold transitions, repeated compaction marks, partial reads, tool/system changes, a branch move, and a one-day synthetic timestamp gap. Correlations are consumed once and never become a verified cause. Restart coverage keeps persisted aggregate usage separate from a fresh live baseline. Existing error/aborted tests cover skipped replies. | +| [`session-metrics.test.ts`](../../tests/extensions/model-info/session-metrics.test.ts) | Existing regressions separately retain active-branch usage from assistant/tool/compaction/branch-summary entries and branch changes. | + +Run from the checkout with the repository-supported Node version: + +```sh +node --test --experimental-strip-types tests/extensions/model-info/cache-diagnostics.test.ts tests/extensions/model-info/index.test.ts tests/extensions/model-info/session-metrics.test.ts +``` + +The focused run passed 39 tests (23 pre-existing, 16 added). Repository-wide `bun run check` and `bun run test` receipts belong to the accompanying PR. The synthetic cases show deterministic contracts only: they measure neither provider false-positive/false-negative rates nor cache cost or latency. + +## Still required for #156 + +No genuine OpenPI / Bare Pi / OMP Session set was collected or replayed here. A future comparison needs independently retrievable, privacy-reviewed evidence with each checkout/package identity, actual loaded OpenPI source (`pi list`), Pi version, provider/API/model identity, task and verifier identity, normalized usage, failures, and relevant lifecycle boundaries. Keep source facts, correlations, and any provider-confirmed cause distinct. A saved Session may not contain the tool/system fingerprints needed to reconstruct every live correlation; absence must remain unknown rather than being filled in from the current environment. + +Preserve raw Sessions outside Git under stable evidence identities and publish only authorized bounded receipts. Do not treat synthetic timestamps as TTL evidence or label an invented trace as a real OpenPI/Bare Pi/OMP run. The real comparative replay, any separately evidenced TTL classification, and the acceptance gate before opt-in UI remain open. There is no new UI, command, model-visible context, provider behavior, or persisted configuration in this contribution. diff --git a/docs/research/README.md b/docs/research/README.md index 25df6042..46cf1891 100644 --- a/docs/research/README.md +++ b/docs/research/README.md @@ -8,6 +8,8 @@ Research records preserve sourced investigation and distinguish observations, in ## Validated investigations +- [`CACHE_USAGE_CONTRACT_2026-09-11.md`](CACHE_USAGE_CONTRACT_2026-09-11.md) — source-scoped Pi usage field matrix and synthetic cache-diagnostic replay boundaries; real comparative Session evidence remains open ([#156](https://github.com/openpi-dev/openpi/issues/156)). + - [`CURSOR_NATIVE_RECOVERY_2026-09-08.md`](CURSOR_NATIVE_RECOVERY_2026-09-08.md) — bounded in-band native execution rejection, Pi-owned tools, cancellation identity, and actual child acceptance ([#234](https://github.com/openpi-dev/openpi/issues/234)). - [`WEB_STARTUP_2026-09-07.md`](WEB_STARTUP_2026-09-07.md) — terminal startup feedback, browser-launch waiting, and exploratory timing limits ([#450](https://github.com/openpi-dev/openpi/issues/450)). diff --git a/tests/extensions/model-info/cache-diagnostics.test.ts b/tests/extensions/model-info/cache-diagnostics.test.ts index bf48e487..33ffbcce 100644 --- a/tests/extensions/model-info/cache-diagnostics.test.ts +++ b/tests/extensions/model-info/cache-diagnostics.test.ts @@ -150,3 +150,113 @@ test("reset removes the prior warm baseline and pending correlations", () => { assert.equal(observation.kind, "first-turn"); assert.deepEqual(observation.correlations, []); }); + +// Synthetic normalized Usage vectors, not captured provider responses. +for (const [provider, semantics] of [ + ["azure-openai-responses", "implicit-best-effort"], + ["google", "implicit-best-effort"], + ["google-antigravity", "implicit-best-effort"], + ["google-gemini-cli", "implicit-best-effort"], + ["openai", "implicit-best-effort"], + ["openai-codex", "implicit-best-effort"], + ["openai-responses", "implicit-best-effort"], + ["amazon-bedrock", "unknown"], + ["google-vertex", "unknown"], + ["openrouter", "unknown"], + ["custom-provider", "unknown"], +] as const) { + test(`${provider} synthetic warm/cold replay never claims an explicit-prefix miss`, () => { + const tracker = createCacheDiagnosticsTracker(); + const observations = [4_096, 0, 0, 4_096, 2_048, 0].map( + (cacheRead, turnIndex) => + tracker.observe({ + turnIndex, + identity: { ...identity, provider }, + usage: usage({ input: 100, cacheRead, cacheWrite: 500 }), + }), + ); + + assert.deepEqual( + observations.map((observation) => observation.kind), + ["first-turn", "unknown", "cold", "warm", "partial-hit", "unknown"], + ); + for (const observation of observations) { + assert.equal(observation.semantics, semantics); + assert.equal(observation.reprocessedTokens, null); + assert.equal(observation.verifiedCause, null); + assert.equal(observation.evidence, "observation"); + } + }); +} + +test("the explicit-prefix warm threshold includes exactly 2048 reported read tokens", () => { + for (const [previousCacheRead, kind] of [ + [2_047, "cold"], + [2_048, "miss-after-warm-prefix"], + ] as const) { + const tracker = createCacheDiagnosticsTracker(); + tracker.observe({ + turnIndex: 0, + identity, + usage: usage({ cacheRead: previousCacheRead }), + }); + const observation = tracker.observe({ + turnIndex: 1, + identity, + usage: usage({ input: 100, cacheWrite: 3_000 }), + }); + + assert.equal(observation.kind, kind); + assert.equal(observation.usage.promptTokens, 3_100); + assert.equal(observation.reprocessedTokens, kind === "cold" ? null : 100); + } +}); + +test("prompt accounting excludes output, reasoning, write subsets, totals and cost", () => { + const reported = usage({ + input: 100, + cacheRead: 2_000, + cacheWrite: 500, + cacheWrite1h: 300, + output: 90, + reasoning: 40, + totalTokens: 2_690, + cost: { input: 1, output: 2, cacheRead: 3, cacheWrite: 4, total: 10 }, + }); + const before = structuredClone(reported); + const observation = createCacheDiagnosticsTracker().observe({ + turnIndex: 0, + identity, + usage: reported, + }); + + assert.deepEqual(observation.usage, { + input: 100, + cacheRead: 2_000, + cacheWrite: 500, + promptTokens: 2_600, + }); + assert.deepEqual(reported, before); +}); + +test("non-finite and negative prompt counters cannot establish a warm baseline", () => { + for (const invalid of [Number.NaN, Number.POSITIVE_INFINITY, -1]) { + const tracker = createCacheDiagnosticsTracker(); + const first = tracker.observe({ + turnIndex: 0, + identity, + usage: usage({ input: invalid, cacheRead: invalid, cacheWrite: invalid }), + }); + assert.deepEqual(first.usage, { + input: 0, + cacheRead: 0, + cacheWrite: 0, + promptTokens: 0, + }); + assert.equal( + tracker.observe({ turnIndex: 1, identity, usage: usage({ input: 100 }) }) + .kind, + "cold", + ); + } +}); diff --git a/tests/extensions/model-info/index.test.ts b/tests/extensions/model-info/index.test.ts index 8648feb4..845be824 100644 --- a/tests/extensions/model-info/index.test.ts +++ b/tests/extensions/model-info/index.test.ts @@ -436,3 +436,156 @@ for (const stopReason of ["error", "aborted"] as const) { assert.deepEqual(observation?.correlations, ["compaction"]); }); } + +test("synthetic event replay consumes correlations once and keeps elapsed time causally unknown", async () => { + // Hand-authored events exercise OpenPI's extension seam. These are not a + // Session recording, a provider response fixture, or a TTL experiment. + const harness = new ModelInfoHarness([]); + await harness.selectModel({ + provider: "anthropic", + id: "synthetic-model", + name: "Synthetic model", + contextWindow: 200_000, + reasoning: false, + }); + await harness.emit("session_start"); + + const turns = [ + { input: 100, cacheRead: 4_096, cacheWrite: 0 }, + { input: 1_000, cacheRead: 0, cacheWrite: 3_000 }, + { input: 4_100, cacheRead: 0, cacheWrite: 0 }, + { input: 100, cacheRead: 8_192, cacheWrite: 0 }, + { input: 100, cacheRead: 4_096, cacheWrite: 0 }, + { input: 4_400, cacheRead: 0, cacheWrite: 0 }, + { input: 4_500, cacheRead: 0, cacheWrite: 0 }, + ]; + for (const [turnIndex, counters] of turns.entries()) { + if (turnIndex === 1) { + await harness.emit("session_compact"); + await harness.emit("session_compact"); + } + if (turnIndex === 5) await harness.emit("session_tree"); + await harness.emit("before_agent_start", { + systemPrompt: turnIndex < 4 ? "synthetic system" : "changed system", + systemPromptOptions: { + cwd: "/synthetic", + selectedTools: turnIndex < 4 ? ["read"] : ["read", "bash"], + }, + }); + await harness.emit("turn_end", { + turnIndex, + message: { + ...assistant(`synthetic-${turnIndex}`, null, usage(counters)).message, + api: "anthropic-messages", + provider: "anthropic", + model: "synthetic-model", + // A one-day timestamp gap is deliberately not evidence of expiry. + timestamp: turnIndex < 5 ? turnIndex * 1_000 : 86_400_000 + turnIndex, + }, + toolResults: [], + }); + } + + assert.deepEqual( + harness.cacheObservations.map((observation) => ({ + turnIndex: observation.turnIndex, + kind: observation.kind, + previousCacheRead: observation.previousCacheRead, + reprocessedTokens: observation.reprocessedTokens, + correlations: observation.correlations, + })), + [ + { + turnIndex: 0, + kind: "first-turn", + previousCacheRead: null, + reprocessedTokens: null, + correlations: [], + }, + { + turnIndex: 1, + kind: "miss-after-warm-prefix", + previousCacheRead: 4_096, + reprocessedTokens: 1_000, + correlations: ["compaction"], + }, + { + turnIndex: 2, + kind: "cold", + previousCacheRead: 0, + reprocessedTokens: null, + correlations: [], + }, + { + turnIndex: 3, + kind: "warm", + previousCacheRead: 0, + reprocessedTokens: null, + correlations: [], + }, + { + turnIndex: 4, + kind: "partial-hit", + previousCacheRead: 8_192, + reprocessedTokens: null, + correlations: ["tool-surface-change", "system-prompt-change"], + }, + { + turnIndex: 5, + kind: "miss-after-warm-prefix", + previousCacheRead: 4_096, + reprocessedTokens: 4_400, + correlations: ["branch-change"], + }, + { + turnIndex: 6, + kind: "cold", + previousCacheRead: 0, + reprocessedTokens: null, + correlations: [], + }, + ], + ); + for (const observation of harness.cacheObservations) { + assert.equal(observation.evidence, "observation"); + assert.equal(observation.verifiedCause, null); + } + assert.deepEqual(harness.cacheObservations[1]?.usage, { + input: 1_000, + cacheRead: 0, + cacheWrite: 3_000, + promptTokens: 4_000, + }); +}); + +test("synthetic Session restart does not replay persisted usage into a live warm baseline", async () => { + const harness = new ModelInfoHarness([ + assistant("persisted-warm", null, usage({ input: 100, cacheRead: 4_096 })), + ]); + const begin = () => + harness.emit("before_agent_start", { + systemPrompt: "synthetic system", + systemPromptOptions: { cwd: "/synthetic", selectedTools: ["read"] }, + }); + await harness.emit("session_start"); + await begin(); + await harness.emit("turn_end", { + turnIndex: 0, + message: assistant("live-warm", null, usage({ cacheRead: 4_096 })).message, + toolResults: [], + }); + await harness.emit("session_tree"); + await harness.emit("session_start"); + assert.equal(harness.state.cachePercent, (4_096 / 4_196) * 100); + + await begin(); + await harness.emit("turn_end", { + turnIndex: 0, + message: assistant("new-cold", null, usage({ input: 4_400 })).message, + toolResults: [], + }); + const observation = harness.cacheObservations.at(-1); + assert.equal(observation?.kind, "first-turn"); + assert.equal(observation?.previousCacheRead, null); + assert.deepEqual(observation?.correlations, []); +});