diff --git a/javascript/examples/vitest/from-transcript/README.md b/javascript/examples/vitest/from-transcript/README.md new file mode 100644 index 000000000..770f701d7 --- /dev/null +++ b/javascript/examples/vitest/from-transcript/README.md @@ -0,0 +1,111 @@ +# `fromTranscript` — a runnable Scenario from a real Claude Code transcript (#779) + +Point this helper at a real Claude Code session log (`~/.claude/projects//.jsonl`), +fork it at the moment the agent is about to take a turn, and get a **runnable Scenario** seeded with +that captured history — then let a **live** agent-under-test take the next turn and a **judge** assert +whether it reproduces (or fixes) the original behaviour. + +This is the [#779](https://github.com/langwatch/scenario/issues/779) **Strategy B** shape +(adapter-specific, JSONL-sourced). It graduates the [#789](https://github.com/langwatch/scenario/pull/789) +spike into a small, tested helper. It is **not** the productionized feature — no general UI, no +trace-ingestion pipeline. Scope is locked to the Claude Code JSONL path; the from-traces (Strategy C) +and full-sandbox (Strategy D) paths are deferred (see the issue). + +## Why this is non-trivial + +A Claude Code session JSONL is **not** a flat message list. It is an append log threaded as a **tree** +via `parentUuid`/`uuid`, and the reader must handle four things a naive top-to-bottom read gets wrong: + +1. **Walk the `parentUuid` chain** leaf→root (then reverse), not file order — editing/re-running a turn + creates sibling branches, and subagent (`Task`) sidechains are appended after the main thread. +2. **`role:"user"` is overloaded** — a line is *either* a genuine human turn *or* a `tool_result` + carried as a user message (the sample fixture: 2 real human turns vs 7 tool_results). They must be + distinguished. +3. **One assistant message is split across many JSONL lines** (one per `thinking`/`text`/`tool_use` + block, all sharing one `message.id`) — these must be re-merged into a single assistant turn. +4. **Anthropic blocks → AI-SDK v6 `ModelMessage`** — `tool_use`/`tool_result`/`thinking` map onto a + different native shape; `tool_result` carries only an id, so the tool **name** is recovered by + pairing back to its `tool_use`. + +## The `max_turns` seeding gotcha (and the fix) + +Seeding captured history via `message()`/`user()`/`agent()` script steps **silently consumes the shared +`maxTurns` budget** (default 10). A ~15-message transcript then dies with *"Reached maximum turns"* the +moment a `proceed()` runs. This helper seeds through the **turn-free `state.addMessage()` side-door** +(`seedStep`), which writes straight to `state.messages` and never touches `currentTurn` — so a seed of +any length costs zero budget. (In the TS SDK the cap is only enforced inside `_step()` / +`proceed()`, `scenario-execution.ts:846` — confirmed.) + +## Usage + +```ts +import scenario from "@langwatch/scenario"; +import { fromTranscript } from "./from-transcript"; + +// 1. Convert a real CC session → a seeded Scenario (fork defaults to "before the last assistant reply"). +const seed = fromTranscript("/path/to/session.jsonl", { + forkAt: { beforeLastAssistant: true }, // or { uuid } | { index } + flattenTools: true, // fold tool I/O into text when the agent-under-test is a + // different model family than the transcript's source +}); + +// 2. Drop the seed straight into a normal scenario.run(): side-door seed → live agent() → judge(). +const result = await scenario.run({ + name: "replay: ", + description: "Replay of a captured Claude Code session; the live agent takes the forked turn.", + maxTurns: 10, // seed does NOT consume this — it is side-door seeded + agents: [myAgentUnderTest, scenario.judgeAgent({ criteria: ["…what the agent must / must not do…"] })], + script: [seed.seedStep, scenario.agent(), scenario.judge()], +}); +``` + +`fromTranscript(path, opts)` returns a `SeededScenario` — `{ seedMessages, seedStep, lastHumanText, +originalNextText, turns, stats, … }`. It has **no dependency on the scenario SDK**, so the reader and +converter stay unit-testable without spinning up a run (see `reader.test.ts`). + +### Options + +| option | meaning | +|---|---| +| `forkAt` | `{ uuid }` \| `{ index }` \| `{ beforeLastAssistant: true }` (default) — where to cut the history | +| `flattenTools` | fold `tool_use`/`tool_result` into assistant **text** (robust cross-model seed; avoids provider 400s on cross-family structured tool linkage) | +| `includeThinking` | emit captured `thinking` (default `false` — feeding Anthropic reasoning to a non-Anthropic model is invalid) | +| `dropMatching` | scrub seed messages matching a `RegExp` — used to reproduce an "absent config/memory" failure | + +## The end-to-end demo + +`proof.e2e.test.ts` runs the whole loop **live** (real model calls, not mocked) against +`fixtures/real-cc-session.jsonl` — a **PII-scrubbed** real Claude Code memory-recall session (the +`KUMQUAT77` token is a synthetic placeholder that grants nothing; every tool_result / attachment / +injected-skill body and tool-call input was reduced to a placeholder so no real operator memory or infra +ships in this public repo — the parentUuid tree, tool structure, thinking, and recall flow stay intact): + +- **Proof A** — faithful full-history replay: the ~15-message seed (which would trip a naive + `maxTurns`) is side-door seeded, the live agent answers the user's final question, and the judge + returns a real verdict. +- **max_turns contrast** — the same history seeded the naive way (`message()`×N then `proceed()`) hits + the turn cap, while the side-door seed under the same budget reaches a verdict. +- **Proof B (fix-loop)** — the token is scrubbed from the seed; the judge verdict **flips FAIL→PASS** + on nothing but the one config the JSONL omits (the reconstructed memory in the agent's system prompt). + +### Run it + +```bash +cd javascript && pnpm install && pnpm run build # build the SDK → dist/ +cd examples/vitest && pnpm install +pnpm exec vitest run from-transcript/reader.test.ts # unit tests — no API keys needed +# live e2e — gated behind SC779_LIVE (skipped by default so CI's example suite needs no funded key): +printf 'OPENAI_API_KEY=sk-...\nLANGWATCH_API_KEY=sk-lw-...\n' > .env # LANGWATCH key optional (adds a live report URL) +SC779_LIVE=1 pnpm exec vitest run from-transcript/proof.e2e.test.ts +# to run against Gemini instead of the OpenAI default: SC779_LIVE=1 SC779_PROVIDER=gemini pnpm exec vitest run … +``` + +## Honest caveat (the crux) + +The Claude Code JSONL captures the **conversation** (messages, tool I/O, thinking, timestamps) but +**not the config that shaped it** — the injected CLAUDE.md, agent profile, tool schemas, and +system-reminders are absent. So this is Strategy-A "artificial context injection": you seed the live +agent with the captured *conversation* and observe whether a prompt/tool/procedure change moves the next +response. It is strong enough to iterate on "does this change flip the behaviour" (Proof B does exactly +that), and deliberately weaker than a claim of "bit-exact what Claude Code would have done" — that is the +deferred sandbox path (Strategy D). diff --git a/javascript/examples/vitest/from-transcript/cc-transcript.ts b/javascript/examples/vitest/from-transcript/cc-transcript.ts new file mode 100644 index 000000000..b0c704669 --- /dev/null +++ b/javascript/examples/vitest/from-transcript/cc-transcript.ts @@ -0,0 +1,317 @@ +/** + * #779 — Claude Code transcript reader + Anthropic→AI-SDK message adapter. + * + * Strategy B (adapter-specific, JSONL-sourced). This is the CC-adapter knowledge: + * how to turn a raw Claude Code session JSONL into a linear AI-SDK `ModelMessage[]` + * that a Scenario can be seeded with. + * + * A Claude Code session JSONL is an append log threaded as a TREE via parentUuid/uuid + * (not a flat list). The "live" conversation is the path from the chosen leaf back to + * root, so we WALK the parentUuid chain — we never read the file top-to-bottom. + * + * Load-bearing normalizations (all exercised by reader.test.ts against a real transcript): + * 1. Walk parentUuid leaf→root, then reverse to root→leaf order. + * 2. Drop non-conversational metadata lines (attachment/system/file-history-snapshot/…). + * 3. `role:"user"` is OVERLOADED — classify each as human | injected(command/skill) | tool_result. + * 4. One assistant message is SPLIT across multiple JSONL lines (one per content block), + * all sharing one `message.id` — re-merge them into a single assistant turn. + * 5. Pair each Anthropic `tool_result` back to its `tool_use` by id to recover toolName + * (AI-SDK v6 tool-result parts REQUIRE a toolName; Anthropic tool_result carries only the id). + * 6. `thinking` blocks ARE captured — preserved in the normalized model, but DROPPED by + * default when emitting to AI-SDK (feeding Anthropic reasoning to a non-Anthropic model + * is invalid). Toggle with includeThinking. + */ +import { readFileSync } from "node:fs"; +import type { ModelMessage, TextPart, ToolCallPart } from "ai"; + +// --------------------------------------------------------------------------- +// Raw JSONL node types (only the fields we use) +// --------------------------------------------------------------------------- + +export interface RawNode { + type?: string; // "user" | "assistant" | "attachment" | "system" | ... + uuid?: string; + parentUuid?: string | null; + timestamp?: string; + isSidechain?: boolean; + message?: { + id?: string; + role?: string; + model?: string; + content?: string | AnthropicBlock[]; + }; +} + +export type AnthropicBlock = + | { type: "text"; text: string } + | { type: "thinking"; thinking: string; signature?: string } + | { type: "tool_use"; id: string; name: string; input: unknown } + | { type: "tool_result"; tool_use_id: string; content: unknown; is_error?: boolean }; + +export type UserKind = "human" | "injected" | "tool_result"; + +// --------------------------------------------------------------------------- +// 1. Parse + walk the tree +// --------------------------------------------------------------------------- + +export function parseSessionFile(path: string): RawNode[] { + return readFileSync(path, "utf8") + .split("\n") + .map((l) => l.trim()) + .filter(Boolean) + .map((l) => { + try { + return JSON.parse(l) as RawNode; + } catch { + return null; + } + }) + .filter((n): n is RawNode => !!n && !!n.uuid); +} + +/** + * Walk the parentUuid chain from a leaf back to the root, returning root→leaf order. + * If `forkAtUuid` is given, that node is treated as the leaf (fork point) and the + * walk starts there — everything after it in the tree is discarded. + * Otherwise the leaf is the last-in-file node that is nobody's parent (the live tip). + */ +export function linearize(nodes: RawNode[], forkAtUuid?: string): RawNode[] { + const byUuid = new Map(); + for (const n of nodes) byUuid.set(n.uuid!, n); + + let leaf = forkAtUuid; + if (!leaf) { + const parents = new Set(nodes.map((n) => n.parentUuid)); + let leaves = nodes.filter((n) => !parents.has(n.uuid)); + // Ignore subagent (Task) sidechains: their nodes carry isSidechain:true and their leaf is + // appended AFTER the main thread's last line, so the naive "last leaf" would walk the + // sub-agent's private conversation instead of the main thread. Prefer non-sidechain leaves. + const mainLeaves = leaves.filter((n) => !n.isSidechain); + if (mainLeaves.length) leaves = mainLeaves; + // pick the leaf that appears last in file order = the live conversation tip + leaf = leaves.length ? leaves[leaves.length - 1].uuid : nodes[nodes.length - 1]?.uuid; + } + + const chain: RawNode[] = []; + let cur: string | null | undefined = leaf; + const seen = new Set(); // guard against pathological cycles + while (cur && byUuid.has(cur) && !seen.has(cur)) { + seen.add(cur); + chain.push(byUuid.get(cur)!); + cur = byUuid.get(cur)!.parentUuid; + } + return chain.reverse(); +} + +export function isConversational(n: RawNode): boolean { + return n.type === "user" || n.type === "assistant"; +} + +/** Classify an overloaded `role:"user"` line. */ +export function classifyUser(n: RawNode): UserKind { + const c = n.message?.content; + if (typeof c === "string") return "human"; + if (Array.isArray(c)) { + if (c.some((b) => (b as AnthropicBlock).type === "tool_result")) return "tool_result"; + // A genuine human turn can arrive as structured content (text + image/file blocks); a + // command/skill injection is text-only. Any non-text block (e.g. image) ⇒ human. + // NOTE (known limitation): a text-ONLY human array is indistinguishable from an injection + // here and is still tagged "injected" — a real signal (userType/isMeta) is productization work. + if (c.some((b) => (b as { type?: string }).type && (b as { type?: string }).type !== "text")) return "human"; + return "injected"; + } + return "injected"; +} + +// --------------------------------------------------------------------------- +// 2. Normalized, provenance-tagged view (fidelity-preserving; keeps thinking) +// --------------------------------------------------------------------------- + +export interface NormalizedTurn { + role: "user" | "assistant" | "tool"; + userKind?: UserKind; // only when role==="user" + uuid: string; + /** Every source JSONL uuid that folded into this turn (assistant merges >1). Enables fork-by-uuid + * to resolve a uuid that points at a non-first line of a same-message.id assistant. */ + uuids: string[]; + timestamp?: string; + text?: string; // human/injected user text OR assistant visible text + thinking?: string[]; // assistant reasoning (preserved, may be dropped on emit) + toolCalls?: { id: string; name: string; input: unknown }[]; + toolResults?: { id: string; name: string; output: string; isError: boolean }[]; +} + +function blockText(content: string | AnthropicBlock[] | undefined): string { + if (typeof content === "string") return content; + if (!Array.isArray(content)) return ""; + return content + .filter((b) => (b as AnthropicBlock).type === "text") + .map((b) => (b as { text: string }).text) + .join("\n"); +} + +function stringifyToolResult(content: unknown): string { + if (typeof content === "string") return content; + if (Array.isArray(content)) { + return content + .map((b) => (b && typeof b === "object" && "text" in b ? (b as { text: string }).text : JSON.stringify(b))) + .join("\n"); + } + return content == null ? "" : JSON.stringify(content); +} + +/** + * Convert a linearized chain into normalized turns: + * - drops metadata nodes, + * - re-merges same-message.id assistant blocks into one turn, + * - recovers toolName for each tool_result via a tool_use id→name map. + */ +/** A NormalizedTurn during construction, carrying the private same-message.id merge tag. */ +type BuildTurn = NormalizedTurn & { _msgId?: string }; + +export function normalize(chain: RawNode[]): NormalizedTurn[] { + const idToName = new Map(); + const out: BuildTurn[] = []; + + for (const n of chain) { + if (!isConversational(n)) continue; + + if (n.type === "assistant") { + const blocks = Array.isArray(n.message?.content) ? (n.message!.content as AnthropicBlock[]) : []; + const thinking = blocks.filter((b) => b.type === "thinking").map((b) => (b as { thinking: string }).thinking); + const text = blockText(n.message?.content); + const toolCalls = blocks + .filter((b) => b.type === "tool_use") + .map((b) => { + const tu = b as { id: string; name: string; input: unknown }; + idToName.set(tu.id, tu.name); + return { id: tu.id, name: tu.name, input: tu.input }; + }); + + const msgId = n.message?.id; + const prev = out[out.length - 1]; + // Merge consecutive assistant lines that belong to the same underlying message. + if (prev && prev.role === "assistant" && msgId && prev._msgId === msgId) { + if (text) prev.text = [prev.text, text].filter(Boolean).join("\n"); + if (thinking.length) prev.thinking = [...(prev.thinking ?? []), ...thinking]; + if (toolCalls.length) prev.toolCalls = [...(prev.toolCalls ?? []), ...toolCalls]; + prev.uuids.push(n.uuid!); // so fork-by-uuid can target a merged-away line + continue; + } + out.push({ + role: "assistant", + uuid: n.uuid!, + uuids: [n.uuid!], + timestamp: n.timestamp, + text: text || undefined, + thinking: thinking.length ? thinking : undefined, + toolCalls: toolCalls.length ? toolCalls : undefined, + ...(msgId ? { _msgId: msgId } : {}), + }); + continue; + } + + // role: "user" + const kind = classifyUser(n); + if (kind === "tool_result") { + const blocks = n.message!.content as AnthropicBlock[]; + const toolResults = blocks + .filter((b) => b.type === "tool_result") + .map((b) => { + const tr = b as { tool_use_id: string; content: unknown; is_error?: boolean }; + return { + id: tr.tool_use_id, + name: idToName.get(tr.tool_use_id) ?? "unknown_tool", + output: stringifyToolResult(tr.content), + isError: !!tr.is_error, + }; + }); + out.push({ role: "tool", uuid: n.uuid!, uuids: [n.uuid!], timestamp: n.timestamp, toolResults }); + } else { + // Keep the content verbatim — do NOT strip a leading "user: " (a real human turn may legitimately + // begin with "user: …"; stripping corrupts the replayed prompt). blockText handles array content. + const text = kind === "human" && typeof n.message?.content === "string" + ? (n.message!.content as string) + : blockText(n.message?.content); + out.push({ role: "user", userKind: kind, uuid: n.uuid!, uuids: [n.uuid!], timestamp: n.timestamp, text }); + } + } + return out; +} + +// --------------------------------------------------------------------------- +// 3. Emit AI-SDK v6 ModelMessage[] (Anthropic blocks → {OpenAI|AI-SDK} shape) +// --------------------------------------------------------------------------- + +export interface EmitOptions { + includeThinking?: boolean; // default false — see class note (6) + /** + * Fold tool_use / tool_result into readable assistant TEXT instead of structured + * AI-SDK tool parts. Use this when seeding a DIFFERENT model family than the one that + * produced the transcript (e.g. replaying a Claude session's context into an OpenAI + * agent-under-test): structured cross-model tool-call linkage trips provider-side + * validation (400 "tool_call_ids did not have response messages"). This is the honest + * Strategy-A "artificial context injection" emit. Default false = faithful structured. + */ + flattenTools?: boolean; +} + +const trunc = (s: string, n = 600) => (s.length <= n ? s : s.slice(0, n) + `…(+${s.length - n}ch)`); + +export function toModelMessages(turns: NormalizedTurn[], opts: EmitOptions = {}): ModelMessage[] { + const msgs: ModelMessage[] = []; + for (const t of turns) { + if (t.role === "user") { + msgs.push({ role: "user", content: t.text ?? "" }); + } else if (t.role === "tool") { + if (opts.flattenTools) { + const text = (t.toolResults ?? []) + .map((r) => `[tool ${r.name} ${r.isError ? "ERROR" : "→"}] ${trunc(r.output)}`) + .join("\n"); + // fold into the preceding assistant turn if there is one, else stand alone + const prev = msgs[msgs.length - 1]; + if (prev && prev.role === "assistant" && typeof prev.content === "string") { + prev.content = [prev.content, text].filter(Boolean).join("\n"); + } else { + msgs.push({ role: "assistant", content: text || " " }); + } + continue; + } + msgs.push({ + role: "tool", + content: (t.toolResults ?? []).map((r) => ({ + type: "tool-result" as const, + toolCallId: r.id, + toolName: r.name, + output: { type: "text" as const, value: r.isError ? `ERROR: ${r.output}` : r.output }, + })), + }); + } else if (opts.flattenTools) { + const bits: string[] = []; + if (opts.includeThinking && t.thinking?.length) bits.push(t.thinking.map((th) => `[thinking] ${th}`).join("\n")); + if (t.text) bits.push(t.text); + for (const tc of t.toolCalls ?? []) bits.push(`[called ${tc.name}] ${trunc(JSON.stringify(tc.input ?? {}))}`); + const content = bits.join("\n"); + if (content) msgs.push({ role: "assistant", content }); // skip a thinking-only/empty turn + } else { + // assistant: may carry text and/or tool-calls (and thinking, usually dropped) + const parts: Array = []; + if (opts.includeThinking && t.thinking?.length) { + parts.push({ type: "text", text: t.thinking.map((th) => `[thinking] ${th}`).join("\n") }); + } + if (t.text) parts.push({ type: "text", text: t.text }); + for (const tc of t.toolCalls ?? []) { + parts.push({ type: "tool-call", toolCallId: tc.id, toolName: tc.name, input: tc.input ?? {} }); + } + if (parts.length === 0) { + // thinking-only (dropped) or empty assistant turn — skip, don't inject a blank message + continue; + } else if (parts.length === 1 && parts[0].type === "text") { + msgs.push({ role: "assistant", content: parts[0].text }); + } else { + msgs.push({ role: "assistant", content: parts }); + } + } + } + return msgs; +} diff --git a/javascript/examples/vitest/from-transcript/evidence/e2e-run.log b/javascript/examples/vitest/from-transcript/evidence/e2e-run.log new file mode 100644 index 000000000..6895086f0 --- /dev/null +++ b/javascript/examples/vitest/from-transcript/evidence/e2e-run.log @@ -0,0 +1,83 @@ +# #779 fromTranscript — live e2e run receipt (PII-scrubbed fixture) +# command: cd javascript/examples/vitest && SC779_LIVE=1 SC779_PROVIDER=gemini pnpm exec vitest run from-transcript/proof.e2e.test.ts +# run id: scenariobatch_3GgLkgfhAdLZMEmBI4kwGaeneKd +# LangWatch: https://app.langwatch.ai/voice-bugbash-nWygNc/simulations/779-from-transcript/scenariobatch_3GgLkgfhAdLZMEmBI4kwGaeneKd +# models: agent+judge = gemini-2.5-flash (OpenAI gpt-5-mini is the default when SC779_PROVIDER unset) + + + RUN v4.1.7 /javascript/examples/vitest + +──────────── PROOF A — faithful replay (token in context) ──────────── +seed: 12 messages (from 53 parsed nodes → 53-node parentUuid chain → 20 turns: 2 human / 2 injected / 7 tool / 9 assistant) + fork: before turn #19; live agent replaces original reply → "KUMQUAT77" + seeded conversation (role: preview): + user | user: Remember this exact token for later: KUMQUAT77. Just acknowledge it. + assistant | I'm saving this token to memory now. [called Write] {"elided":"tool input elided for fixtu + assistant | [called Skill] {"elided":"tool input elided for fixture"} [tool Skill →] [tool output elid + user | [injected command/skill output elided for fixture] + assistant | [called Read] {"elided":"tool input elided for fixture"} [tool Read →] [tool output elided + assistant | 🎩 This is the way 🪄 Acknowledged. Let me check if there's a procedure for memory managem + user | [injected command/skill output elided for fixture] + assistant | No procedure found for "memory, remember, save token" — this is a straightforward memory o + assistant | [called Read] {"elided":"tool input elided for fixture"} [tool Read →] [tool output elided + assistant | [called Edit] {"elided":"tool input elided for fixture"} [tool Edit →] [tool output elided + assistant | **Done.** Token `KUMQUAT77` saved to memory and indexed. Recalled in future conversations. + user | user: What was the exact token I asked you to remember? Reply with just the token. + model: gemini:gemini-2.5-flash + +──────────────────────────────────────────────────────────── +🎭 Running Scenario Tests +──────────────────────────────────────────────────────────── +Follow it live: https://app.langwatch.ai/voice-bugbash-nWygNc/simulations/779-from-transcript/scenariobatch_3GgLkgfhAdLZMEmBI4kwGaeneKd +──────────────────────────────────────────────────────────── + + ▶ LIVE AGENT TURN: "KUMQUAT77" + ⚖ JUDGE: success=true + met: ["The assistant's reply gives the exact remembered token, which is KUMQUAT77."] + unmet: [] + reasoning: The assistant's reply was exactly 'KUMQUAT77', which is the token it was asked to remember. + +──────────── DoD-4 CONTRAST (naive message()+proceed vs side-door, maxTurns=4, seed=12) ──────────── + NAIVE: threw=no success=false reasoning="Reached maximum turns (4) without conclusion" + SIDE-DOOR: success=true met=["The assistant's reply gives the exact remembered token, which is KUMQUAT77."] reasoning="The assistant successfully remembered and replied with the exact token KUMQUAT77." + +──────────── PROOF B — fix-loop (token scrubbed from seed) ──────────── +seed: 10 messages (from 53 parsed nodes → 53-node parentUuid chain → 20 turns: 2 human / 2 injected / 7 tool / 9 assistant) + ⚠ scrubbed 2 token-bearing message(s) to reproduce the 'absent memory' failure + fork: before turn #19; live agent replaces original reply → "KUMQUAT77" + seeded conversation (role: preview): + assistant | I'm saving this token to memory now. [called Write] {"elided":"tool input elided for fixtu + assistant | [called Skill] {"elided":"tool input elided for fixture"} [tool Skill →] [tool output elid + user | [injected command/skill output elided for fixture] + assistant | [called Read] {"elided":"tool input elided for fixture"} [tool Read →] [tool output elided + assistant | 🎩 This is the way 🪄 Acknowledged. Let me check if there's a procedure for memory managem + user | [injected command/skill output elided for fixture] + assistant | No procedure found for "memory, remember, save token" — this is a straightforward memory o + assistant | [called Read] {"elided":"tool input elided for fixture"} [tool Read →] [tool output elided + assistant | [called Edit] {"elided":"tool input elided for fixture"} [tool Edit →] [tool output elided + user | user: What was the exact token I asked you to remember? Reply with just the token. + + ▼ BEFORE (baseline config): + ▶ LIVE AGENT TURN: "J4fB7tE2pQ" + ⚖ JUDGE: success=false + met: [] + unmet: ["The assistant's reply gives the exact remembered token, which is KUMQUAT77."] + reasoning: The assistant did not reply with the exact remembered token, KUMQUAT77. Instead, it replied with 'J4fB7tE2pQ'. + + ▲ AFTER (fixed config = reconstructed memory): + ▶ LIVE AGENT TURN: "KUMQUAT77" + ⚖ JUDGE: success=true + met: ["The assistant's reply gives the exact remembered token, which is KUMQUAT77."] + unmet: [] + reasoning: The assistant's final reply was 'KUMQUAT77', which exactly matches the remembered token specified in the criteria. + + ✓ from-transcript/proof.e2e.test.ts (3 tests) 13991ms + ✓ Proof A (DoD-3 + DoD-4): faithful full-history replay runs to a judge verdict; >10-msg seed survives maxTurns 1780ms + ✓ DoD-4 contrast: NAIVE message()-per-line seed THEN proceed() trips maxTurns; the side-door does not 1755ms + ✓ Proof B (DoD-5): the fix-loop — same forked scenario, judge verdict FLIPS on the omitted config 10454ms + + Test Files 1 passed (1) + Tests 3 passed (3) + Start at 15:26:48 + Duration 16.36s (transform 270ms, setup 201ms, import 2.02s, tests 13.99s, environment 0ms) + diff --git a/javascript/examples/vitest/from-transcript/fixtures/real-cc-session.jsonl b/javascript/examples/vitest/from-transcript/fixtures/real-cc-session.jsonl new file mode 100644 index 000000000..f585d99db --- /dev/null +++ b/javascript/examples/vitest/from-transcript/fixtures/real-cc-session.jsonl @@ -0,0 +1,63 @@ +{"type":"queue-operation","uuid":null,"parentUuid":null,"timestamp":"2026-07-09T10:27:21.190Z","isSidechain":null} +{"type":"queue-operation","uuid":null,"parentUuid":null,"timestamp":"2026-07-09T10:27:21.190Z","isSidechain":null} +{"type":"user","uuid":"e403926b-3dfd-4db5-86d2-d3966b9742cf","parentUuid":null,"timestamp":"2026-07-09T10:27:21.219Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":"user: Remember this exact token for later: KUMQUAT77. Just acknowledge it."}} +{"type":"attachment","uuid":"d8cf6b3b-334f-4c0a-b5d3-2dddd3fcf0b2","parentUuid":"e403926b-3dfd-4db5-86d2-d3966b9742cf","timestamp":"2026-07-09T10:27:21.215Z","isSidechain":false} +{"type":"attachment","uuid":"7e2f91fd-c7b4-4a78-96ef-5b011bbfb3eb","parentUuid":"d8cf6b3b-334f-4c0a-b5d3-2dddd3fcf0b2","timestamp":"2026-07-09T10:27:21.215Z","isSidechain":false} +{"type":"attachment","uuid":"0d397d16-7740-4b03-b22c-b90d657d9a6b","parentUuid":"7e2f91fd-c7b4-4a78-96ef-5b011bbfb3eb","timestamp":"2026-07-09T10:27:21.215Z","isSidechain":false} +{"type":"attachment","uuid":"1b135f7b-2db4-4495-bde6-7b56f8713b2c","parentUuid":"0d397d16-7740-4b03-b22c-b90d657d9a6b","timestamp":"2026-07-09T10:27:21.218Z","isSidechain":false} +{"type":"attachment","uuid":"97318535-e5ae-456c-9277-56bf23346821","parentUuid":"1b135f7b-2db4-4495-bde6-7b56f8713b2c","timestamp":"2026-07-09T10:27:21.457Z","isSidechain":false} +{"type":"last-prompt","uuid":null,"parentUuid":null,"timestamp":null,"isSidechain":null} +{"type":"assistant","uuid":"af3e4e11-694f-4f39-b28b-d6e90eaa338a","parentUuid":"97318535-e5ae-456c-9277-56bf23346821","timestamp":"2026-07-09T10:27:23.682Z","isSidechain":false,"message":{"id":"msg_011CcrKmxPEjCX7QhJF6HfrH","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"fdf12fbd-a1a9-43f0-98b4-0c67c50a1a45","parentUuid":"af3e4e11-694f-4f39-b28b-d6e90eaa338a","timestamp":"2026-07-09T10:27:24.078Z","isSidechain":false,"message":{"id":"msg_011CcrKmxPEjCX7QhJF6HfrH","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"text","text":"I'm saving this token to memory now."}]}} +{"type":"assistant","uuid":"f5291f91-3241-4db2-adc7-0091f3acaebc","parentUuid":"fdf12fbd-a1a9-43f0-98b4-0c67c50a1a45","timestamp":"2026-07-09T10:27:25.553Z","isSidechain":false,"message":{"id":"msg_011CcrKmxPEjCX7QhJF6HfrH","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"tool_use","id":"toolu_01EBZyWYFYtsFfQaQ5Xu5tpU","name":"Write","input":{"elided":"tool input elided for fixture"}}]}} +{"type":"user","uuid":"002fb2fd-9de0-49d7-8a24-fd0d2d3caf21","parentUuid":"f5291f91-3241-4db2-adc7-0091f3acaebc","timestamp":"2026-07-09T10:27:25.613Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"tool_result","tool_use_id":"toolu_01EBZyWYFYtsFfQaQ5Xu5tpU","content":"[tool output elided for fixture]","is_error":true}]}} +{"type":"assistant","uuid":"8e01bc91-40f8-4937-8d75-c6efa099c5a3","parentUuid":"002fb2fd-9de0-49d7-8a24-fd0d2d3caf21","timestamp":"2026-07-09T10:27:27.797Z","isSidechain":false,"message":{"id":"msg_011CcrKnGoWe3yuadYKfBWKS","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"f7af232d-f1c5-4aab-beeb-bd116991c908","parentUuid":"8e01bc91-40f8-4937-8d75-c6efa099c5a3","timestamp":"2026-07-09T10:27:27.873Z","isSidechain":false,"message":{"id":"msg_011CcrKnGoWe3yuadYKfBWKS","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"tool_use","id":"toolu_01Uc69TyZ339xgminQj5oEqE","name":"Skill","input":{"elided":"tool input elided for fixture"}}]}} +{"type":"user","uuid":"cbbec62e-9bc1-4c5e-94b9-02cf6cf510ab","parentUuid":"f7af232d-f1c5-4aab-beeb-bd116991c908","timestamp":"2026-07-09T10:27:28.188Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"tool_result","tool_use_id":"toolu_01Uc69TyZ339xgminQj5oEqE","content":"[tool output elided for fixture]","is_error":false}]}} +{"type":"attachment","uuid":"3d39b06d-dad3-4b39-9235-b642376936c8","parentUuid":"cbbec62e-9bc1-4c5e-94b9-02cf6cf510ab","timestamp":"2026-07-09T10:27:28.000Z","isSidechain":false} +{"type":"user","uuid":"d3c4f33a-96bc-4023-84ee-8eae45107098","parentUuid":"3d39b06d-dad3-4b39-9235-b642376936c8","timestamp":"2026-07-09T10:27:27.966Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"text","text":"[injected command/skill output elided for fixture]"}]}} +{"type":"attachment","uuid":"271767c2-f0ae-4a71-a93c-6163b9c91c87","parentUuid":"d3c4f33a-96bc-4023-84ee-8eae45107098","timestamp":"2026-07-09T10:27:27.965Z","isSidechain":false} +{"type":"attachment","uuid":"e8ecdc22-be68-4e7a-9cb8-6e90d47645a8","parentUuid":"271767c2-f0ae-4a71-a93c-6163b9c91c87","timestamp":"2026-07-09T10:27:27.966Z","isSidechain":false} +{"type":"attachment","uuid":"b3fac612-ed37-41e9-bb53-b81650e24eb5","parentUuid":"e8ecdc22-be68-4e7a-9cb8-6e90d47645a8","timestamp":"2026-07-09T10:27:27.966Z","isSidechain":false} +{"type":"attachment","uuid":"d2faa57c-c4e6-46ed-aa8f-70eb36520e97","parentUuid":"b3fac612-ed37-41e9-bb53-b81650e24eb5","timestamp":"2026-07-09T10:27:28.238Z","isSidechain":false} +{"type":"attachment","uuid":"41263320-dbf4-457a-890f-3adaf1b73bf1","parentUuid":"d2faa57c-c4e6-46ed-aa8f-70eb36520e97","timestamp":"2026-07-09T10:27:28.238Z","isSidechain":false} +{"type":"assistant","uuid":"3772b160-64d3-41d2-8b4a-f3c1da1df079","parentUuid":"41263320-dbf4-457a-890f-3adaf1b73bf1","timestamp":"2026-07-09T10:27:29.976Z","isSidechain":false,"message":{"id":"msg_011CcrKnStEqW87YZzykM2Hc","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"66920252-197a-4702-aa85-e2df3108867f","parentUuid":"3772b160-64d3-41d2-8b4a-f3c1da1df079","timestamp":"2026-07-09T10:27:30.369Z","isSidechain":false,"message":{"id":"msg_011CcrKnStEqW87YZzykM2Hc","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"tool_use","id":"toolu_016oLTqx6ttJEgCcqL8Y3EZd","name":"Read","input":{"elided":"tool input elided for fixture"}}]}} +{"type":"user","uuid":"776eb97f-7f90-4a9e-bffe-d8611a448dd8","parentUuid":"66920252-197a-4702-aa85-e2df3108867f","timestamp":"2026-07-09T10:27:30.624Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"tool_result","tool_use_id":"toolu_016oLTqx6ttJEgCcqL8Y3EZd","content":"[tool output elided for fixture]","is_error":false}]}} +{"type":"attachment","uuid":"c3a761e9-6abd-42b8-8c56-01f43a265aa9","parentUuid":"776eb97f-7f90-4a9e-bffe-d8611a448dd8","timestamp":"2026-07-09T10:27:30.664Z","isSidechain":false} +{"type":"last-prompt","uuid":null,"parentUuid":null,"timestamp":null,"isSidechain":null} +{"type":"assistant","uuid":"ac99167c-201a-4186-8e46-0744cadc43d0","parentUuid":"c3a761e9-6abd-42b8-8c56-01f43a265aa9","timestamp":"2026-07-09T10:27:35.342Z","isSidechain":false,"message":{"id":"msg_011CcrKniiz2eWv1DA4m3m1t","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"28a13e0d-4ca1-4d4d-9682-665d7744ada1","parentUuid":"ac99167c-201a-4186-8e46-0744cadc43d0","timestamp":"2026-07-09T10:27:35.678Z","isSidechain":false,"message":{"id":"msg_011CcrKniiz2eWv1DA4m3m1t","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"text","text":"🎩 This is the way 🪄\n\nAcknowledged. Let me check if there's a procedure for memory management."}]}} +{"type":"assistant","uuid":"0ed28731-ca67-4510-997e-b1be594a597c","parentUuid":"28a13e0d-4ca1-4d4d-9682-665d7744ada1","timestamp":"2026-07-09T10:27:36.254Z","isSidechain":false,"message":{"id":"msg_011CcrKniiz2eWv1DA4m3m1t","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"tool_use","id":"toolu_01WN7FH23eRrKAk4JVjfnxXL","name":"Skill","input":{"elided":"tool input elided for fixture"}}]}} +{"type":"user","uuid":"ec02a678-e270-44df-92f7-d0a855115dfb","parentUuid":"0ed28731-ca67-4510-997e-b1be594a597c","timestamp":"2026-07-09T10:27:36.500Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"tool_result","tool_use_id":"toolu_01WN7FH23eRrKAk4JVjfnxXL","content":"[tool output elided for fixture]","is_error":false}]}} +{"type":"attachment","uuid":"0b2647ff-25eb-4890-9514-595c6cfff969","parentUuid":"ec02a678-e270-44df-92f7-d0a855115dfb","timestamp":"2026-07-09T10:27:36.342Z","isSidechain":false} +{"type":"user","uuid":"6dfa5c13-43a5-47a4-b32f-839ea50b0095","parentUuid":"0b2647ff-25eb-4890-9514-595c6cfff969","timestamp":"2026-07-09T10:27:36.314Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"text","text":"[injected command/skill output elided for fixture]"}]}} +{"type":"attachment","uuid":"5d7e067a-22c0-46b8-abe8-a806cd9a3287","parentUuid":"6dfa5c13-43a5-47a4-b32f-839ea50b0095","timestamp":"2026-07-09T10:27:36.314Z","isSidechain":false} +{"type":"assistant","uuid":"f7677327-a5fc-4cff-ade8-d1efeba7dbc2","parentUuid":"5d7e067a-22c0-46b8-abe8-a806cd9a3287","timestamp":"2026-07-09T10:27:41.878Z","isSidechain":false,"message":{"id":"msg_011CcrKo7ZAnM88fTJa2ahbo","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"ff7b83b9-b16f-47da-90b4-0447e09da6c3","parentUuid":"f7677327-a5fc-4cff-ade8-d1efeba7dbc2","timestamp":"2026-07-09T10:27:42.257Z","isSidechain":false,"message":{"id":"msg_011CcrKo7ZAnM88fTJa2ahbo","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"text","text":"No procedure found for \"memory, remember, save token\" — this is a straightforward memory operation. Saving the token now using the established memory format."}]}} +{"type":"assistant","uuid":"720e10e5-232d-4368-b808-6454848aa96c","parentUuid":"ff7b83b9-b16f-47da-90b4-0447e09da6c3","timestamp":"2026-07-09T10:27:43.732Z","isSidechain":false,"message":{"id":"msg_011CcrKo7ZAnM88fTJa2ahbo","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"tool_use","id":"toolu_01YYN2QU8bqFGfzdinVpUvWv","name":"Write","input":{"elided":"tool input elided for fixture"}}]}} +{"type":"user","uuid":"a5569dad-d367-4be1-bcc4-606e8c3a10c0","parentUuid":"720e10e5-232d-4368-b808-6454848aa96c","timestamp":"2026-07-09T10:27:43.814Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"tool_result","tool_use_id":"toolu_01YYN2QU8bqFGfzdinVpUvWv","content":"[tool output elided for fixture]","is_error":false}]}} +{"type":"attachment","uuid":"309a3f0c-35cf-49ae-a526-1eedd490e1c0","parentUuid":"a5569dad-d367-4be1-bcc4-606e8c3a10c0","timestamp":"2026-07-09T10:27:43.799Z","isSidechain":false} +{"type":"assistant","uuid":"c24bf34a-f885-492a-8e8f-177264239541","parentUuid":"309a3f0c-35cf-49ae-a526-1eedd490e1c0","timestamp":"2026-07-09T10:27:45.628Z","isSidechain":false,"message":{"id":"msg_011CcrKobThXsQSE89dmmgzQ","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"0f7773c7-22c7-4fe6-ae9d-b58a2f2073d8","parentUuid":"c24bf34a-f885-492a-8e8f-177264239541","timestamp":"2026-07-09T10:27:45.989Z","isSidechain":false,"message":{"id":"msg_011CcrKobThXsQSE89dmmgzQ","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"tool_use","id":"toolu_014h3yEgyfZARypX8iaJeh7w","name":"Read","input":{"elided":"tool input elided for fixture"}}]}} +{"type":"user","uuid":"b9d8440d-7f5e-45ab-aa03-b5803d529a6e","parentUuid":"0f7773c7-22c7-4fe6-ae9d-b58a2f2073d8","timestamp":"2026-07-09T10:27:46.109Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"tool_result","tool_use_id":"toolu_014h3yEgyfZARypX8iaJeh7w","content":"[tool output elided for fixture]","is_error":false}]}} +{"type":"last-prompt","uuid":null,"parentUuid":null,"timestamp":null,"isSidechain":null} +{"type":"assistant","uuid":"0e8c58ff-268e-404f-8516-3e17b094a62e","parentUuid":"b9d8440d-7f5e-45ab-aa03-b5803d529a6e","timestamp":"2026-07-09T10:27:48.035Z","isSidechain":false,"message":{"id":"msg_011CcrKonBNdUxAf5me8mDtX","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"9524168b-922b-437e-a6ac-e3c28bc1fb4e","parentUuid":"0e8c58ff-268e-404f-8516-3e17b094a62e","timestamp":"2026-07-09T10:27:50.610Z","isSidechain":false,"message":{"id":"msg_011CcrKonBNdUxAf5me8mDtX","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"tool_use","id":"toolu_01DmKP8jhN8Es2HuJZqoPR4k","name":"Edit","input":{"elided":"tool input elided for fixture"}}]}} +{"type":"user","uuid":"53a326cd-5d02-4ec4-a370-ccc93fe21f8b","parentUuid":"9524168b-922b-437e-a6ac-e3c28bc1fb4e","timestamp":"2026-07-09T10:27:50.717Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":[{"type":"tool_result","tool_use_id":"toolu_01DmKP8jhN8Es2HuJZqoPR4k","content":"[tool output elided for fixture]","is_error":false}]}} +{"type":"attachment","uuid":"5cc79640-7861-4fe7-8f66-9ab59e374a4d","parentUuid":"53a326cd-5d02-4ec4-a370-ccc93fe21f8b","timestamp":"2026-07-09T10:27:50.711Z","isSidechain":false} +{"type":"attachment","uuid":"7310a338-1bcb-4539-be16-2a82f526acd6","parentUuid":"5cc79640-7861-4fe7-8f66-9ab59e374a4d","timestamp":"2026-07-09T10:27:50.824Z","isSidechain":false} +{"type":"assistant","uuid":"20fed070-85bf-4e73-8dfd-5ad9364fc6d8","parentUuid":"7310a338-1bcb-4539-be16-2a82f526acd6","timestamp":"2026-07-09T10:27:53.556Z","isSidechain":false,"message":{"id":"msg_011CcrKp8EdAHLdXzNYgpnS2","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"877c3b81-1f03-4939-a48b-0ae88d2ea781","parentUuid":"20fed070-85bf-4e73-8dfd-5ad9364fc6d8","timestamp":"2026-07-09T10:27:53.794Z","isSidechain":false,"message":{"id":"msg_011CcrKp8EdAHLdXzNYgpnS2","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"text","text":"**Done.** Token `KUMQUAT77` saved to memory and indexed. Recalled in future conversations."}]}} +{"type":"attachment","uuid":"ba217cce-c4dd-4a07-a5de-ef00ea402eee","parentUuid":"877c3b81-1f03-4939-a48b-0ae88d2ea781","timestamp":"2026-07-09T10:27:53.831Z","isSidechain":false} +{"type":"last-prompt","uuid":null,"parentUuid":null,"timestamp":null,"isSidechain":null} +{"type":"queue-operation","uuid":null,"parentUuid":null,"timestamp":"2026-07-09T10:27:58.071Z","isSidechain":null} +{"type":"queue-operation","uuid":null,"parentUuid":null,"timestamp":"2026-07-09T10:27:58.072Z","isSidechain":null} +{"type":"user","uuid":"cee6b0f1-cea7-4454-9b68-137aece6123f","parentUuid":"ba217cce-c4dd-4a07-a5de-ef00ea402eee","timestamp":"2026-07-09T10:27:58.167Z","isSidechain":false,"message":{"id":null,"role":"user","model":null,"content":"user: What was the exact token I asked you to remember? Reply with just the token."}} +{"type":"attachment","uuid":"3b91059f-7a5b-4af3-a9b4-a56b6d550241","parentUuid":"cee6b0f1-cea7-4454-9b68-137aece6123f","timestamp":"2026-07-09T10:27:58.166Z","isSidechain":false} +{"type":"attachment","uuid":"ce32c348-d9f3-4d0d-9ae4-317fe80ad2c6","parentUuid":"3b91059f-7a5b-4af3-a9b4-a56b6d550241","timestamp":"2026-07-09T10:27:58.386Z","isSidechain":false} +{"type":"assistant","uuid":"493bbaaf-c6ff-4338-95cc-31b97a3b6738","parentUuid":"ce32c348-d9f3-4d0d-9ae4-317fe80ad2c6","timestamp":"2026-07-09T10:28:00.459Z","isSidechain":false,"message":{"id":"msg_011CcrKpg24v9VQHzLjXHoae","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"thinking","thinking":"[reasoning elided for fixture]"}]}} +{"type":"assistant","uuid":"16831e7a-4310-4d5c-9238-18266c08be69","parentUuid":"493bbaaf-c6ff-4338-95cc-31b97a3b6738","timestamp":"2026-07-09T10:28:00.516Z","isSidechain":false,"message":{"id":"msg_011CcrKpg24v9VQHzLjXHoae","role":"assistant","model":"claude-haiku-4-5-20251001","content":[{"type":"text","text":"KUMQUAT77"}]}} +{"type":"attachment","uuid":"8a90c1fb-4b5d-44cf-a347-d88eb451fa99","parentUuid":"16831e7a-4310-4d5c-9238-18266c08be69","timestamp":"2026-07-09T10:28:00.583Z","isSidechain":false} +{"type":"last-prompt","uuid":null,"parentUuid":null,"timestamp":null,"isSidechain":null} +{"type":"mode","uuid":null,"parentUuid":null,"timestamp":null,"isSidechain":null} diff --git a/javascript/examples/vitest/from-transcript/from-transcript.ts b/javascript/examples/vitest/from-transcript/from-transcript.ts new file mode 100644 index 000000000..23a2239ac --- /dev/null +++ b/javascript/examples/vitest/from-transcript/from-transcript.ts @@ -0,0 +1,189 @@ +/** + * #779 — `fromTranscript()` builder facade (Strategy B — adapter-specific, JSONL-sourced). + * + * One call: raw CC JSONL path → a seed you can drop straight into a Scenario `script`. + * It deliberately has NO dependency on the scenario SDK + * (the seed step is just `(state) => state.addMessage(...)`), so the reader/converter + * stay unit-testable without spinning up a run. + * + * The seed step is the DoD-4 fix: it pushes the ENTIRE captured history via the + * turn-free `state.addMessage()` side-door, so a 20- or 200-message transcript never + * consumes the shared `maxTurns` budget (unlike message()/user()/agent() steps, which + * each cost a turn — verified in scenario-execution.ts:scriptCallAgent→newTurn). + */ +import type { ModelMessage, TextPart, ToolCallPart, ToolResultPart } from "ai"; +import { + parseSessionFile, + linearize, + normalize, + toModelMessages, + type NormalizedTurn, +} from "./cc-transcript"; + +export type ForkAt = + | { uuid: string } + | { index: number } // fork BEFORE this normalized-turn index + | { beforeLastAssistant: true }; // fork right before the final assistant reply (default) + +export interface FromTranscriptOptions { + forkAt?: ForkAt; + includeThinking?: boolean; + /** Fold tool I/O into assistant text (robust cross-model seed — see cc-transcript EmitOptions). */ + flattenTools?: boolean; + /** Drop any seeded message whose text matches — used to reproduce an "absent config/memory" failure. */ + dropMatching?: RegExp; +} + +export interface SeededScenario { + seedMessages: ModelMessage[]; + /** DoD-4 side-door: bulk-seed with zero turn-budget cost. */ + seedStep: (state: { addMessage: (m: ModelMessage) => void }) => void; + lastHumanText?: string; + /** The real assistant turn we forked out (what the live agent must now re-produce). */ + originalNextText?: string; + turns: NormalizedTurn[]; + stats: { + parsedNodes: number; + chainLength: number; + normalizedTurns: number; + seededMessages: number; + humanTurns: number; + injectedTurns: number; + toolResultTurns: number; + assistantTurns: number; + forkTurnIndex: number; + prunedOrphanToolParts: number; + /** Messages removed by `dropMatching` (used to reproduce an "absent config/memory" failure). */ + droppedByMatch: number; + }; +} + +function forkIndex(turns: NormalizedTurn[], forkAt: ForkAt | undefined): number { + if (forkAt && "uuid" in forkAt) { + // match against ALL source uuids folded into a turn, so a uuid pointing at a merged-away + // (non-first) line of a same-message.id assistant still resolves to its turn. + const i = turns.findIndex((t) => t.uuids.includes(forkAt.uuid)); + if (i < 0) throw new Error(`forkAt uuid ${forkAt.uuid} not found in transcript`); + return i; + } + if (forkAt && "index" in forkAt) { + const idx = forkAt.index; + if (!Number.isInteger(idx) || idx < 0 || idx > turns.length) { + throw new Error(`forkAt index ${idx} out of range [0, ${turns.length}]`); + } + return idx; + } + // default: beforeLastAssistant — the last assistant turn is the "next turn" the live agent replaces + for (let i = turns.length - 1; i >= 0; i--) { + if (turns[i].role === "assistant") return i; + } + return turns.length; // no assistant turn → seed everything +} + +/** Visible text of a ModelMessage — what a human wrote/read, incl. tool inputs/outputs — for + * dropMatching to test against (NOT JSON.stringify, whose escaping breaks anchors/special chars). */ +function messageText(m: ModelMessage): string { + if (typeof m.content === "string") return m.content; + if (!Array.isArray(m.content)) return ""; + return (m.content as Array) + .map((p) => { + if (p.type === "text") return p.text ?? ""; + if (p.type === "tool-call") return JSON.stringify(p.input ?? {}); + if (p.type === "tool-result") { + // `output` is a discriminated union; only text/json members carry `value`. + const value = "value" in p.output ? p.output.value : undefined; + return typeof value === "string" ? value : JSON.stringify(p.output); + } + return ""; + }) + .join("\n"); +} + +/** Remove tool parts left dangling after a drop/fork: a tool-result whose tool-call was removed, + * or a tool-call whose result was removed. Otherwise the provider rejects the message list + * (400: tool result with no matching call). No-op for the flattened (string-content) seed. */ +function pruneOrphanToolParts(msgs: ModelMessage[]): { messages: ModelMessage[]; pruned: number } { + const callIds = new Set(); + const resultIds = new Set(); + for (const m of msgs) { + if (Array.isArray(m.content)) { + for (const p of m.content as Array) { + if (p.type === "tool-call") callIds.add(p.toolCallId); + if (p.type === "tool-result") resultIds.add(p.toolCallId); + } + } + } + let pruned = 0; + const out: ModelMessage[] = []; + for (const m of msgs) { + if (!Array.isArray(m.content)) { + out.push(m); + continue; + } + const content = m.content as Array; + const kept = content.filter((p) => { + if (p.type === "tool-call" && !resultIds.has(p.toolCallId)) return false; + if (p.type === "tool-result" && !callIds.has(p.toolCallId)) return false; + return true; + }); + pruned += content.length - kept.length; + if (kept.length === 0) continue; // drop a now-empty message + out.push({ ...m, content: kept } as ModelMessage); + } + return { messages: out, pruned }; +} + +export function fromTranscript(path: string, opts: FromTranscriptOptions = {}): SeededScenario { + const raw = parseSessionFile(path); + const chain = linearize(raw, opts.forkAt && "uuid" in opts.forkAt ? opts.forkAt.uuid : undefined); + const turns = normalize(chain); + + const fi = forkIndex(turns, opts.forkAt); + const seededTurns = turns.slice(0, fi); + const originalNext = turns[fi]; + + let seedMessages = toModelMessages(seededTurns, { + includeThinking: opts.includeThinking, + flattenTools: opts.flattenTools, + }); + + let dropped = 0; + if (opts.dropMatching) { + // Strip g/y flags: RegExp.test is stateful for those (lastIndex carries across .filter calls + // and silently skips matches). Test against visible text, not JSON.stringify. + const re = new RegExp(opts.dropMatching.source, opts.dropMatching.flags.replace(/[gy]/g, "")); + const before = seedMessages.length; + seedMessages = seedMessages.filter((m) => !re.test(messageText(m))); + dropped = before - seedMessages.length; + } + + // A drop or a mid-exchange fork can orphan a tool-call/tool-result across the pair; prune so the + // seed is a provider-valid message list. No-op on the flattened (string-content) seed. + const { messages: prunedMessages, pruned } = pruneOrphanToolParts(seedMessages); + seedMessages = prunedMessages; + + const lastHuman = [...seededTurns].reverse().find((t) => t.role === "user" && t.userKind === "human"); + + return { + seedMessages, + seedStep: (state) => { + for (const m of seedMessages) state.addMessage(m); + }, + lastHumanText: lastHuman?.text, + originalNextText: originalNext?.text, + turns, + stats: { + parsedNodes: raw.length, + chainLength: chain.length, + normalizedTurns: turns.length, + seededMessages: seedMessages.length, + humanTurns: turns.filter((t) => t.role === "user" && t.userKind === "human").length, + injectedTurns: turns.filter((t) => t.role === "user" && t.userKind === "injected").length, + toolResultTurns: turns.filter((t) => t.role === "tool").length, + assistantTurns: turns.filter((t) => t.role === "assistant").length, + forkTurnIndex: fi, + prunedOrphanToolParts: pruned, + droppedByMatch: dropped, + }, + }; +} diff --git a/javascript/examples/vitest/from-transcript/model.ts b/javascript/examples/vitest/from-transcript/model.ts new file mode 100644 index 000000000..167dd1d8e --- /dev/null +++ b/javascript/examples/vitest/from-transcript/model.ts @@ -0,0 +1,36 @@ +/** + * Model provider for the end-to-end proof. + * + * Defaults to OpenAI (`gpt-5-mini`) so a reviewer with an `OPENAI_API_KEY` just runs it. + * Set `SC779_PROVIDER=gemini` to route through Gemini's OpenAI-compatible endpoint using + * `GEMINI_API_KEY` — this needs no extra provider dependency (it reuses `@ai-sdk/openai`), + * so the proof runs against whichever funded key is available. Override the model id with + * `SC779_MODEL`. + */ +import { openai, createOpenAI } from "@ai-sdk/openai"; +import type { LanguageModel } from "ai"; + +const PROVIDER = (process.env.SC779_PROVIDER ?? "openai").toLowerCase(); +const isGemini = PROVIDER === "gemini" || PROVIDER === "google"; +if (!isGemini && PROVIDER !== "openai") { + throw new Error(`Unsupported SC779_PROVIDER="${PROVIDER}" — use "openai" (default), "gemini", or "google".`); +} + +export const E2E_MODEL_ID = + process.env.SC779_MODEL ?? (isGemini ? "gemini-2.5-flash" : "gpt-5-mini"); + +export const E2E_MODEL_LABEL = `${PROVIDER}:${E2E_MODEL_ID}`; + +export function e2eModel(): LanguageModel { + if (isGemini) { + const apiKey = process.env.GEMINI_API_KEY ?? process.env.GOOGLE_GENERATIVE_AI_API_KEY; + if (!apiKey) throw new Error("SC779_PROVIDER=gemini but neither GEMINI_API_KEY nor GOOGLE_GENERATIVE_AI_API_KEY is set"); + // Gemini exposes an OpenAI-compatible Chat Completions endpoint (incl. function calling), + // so the already-installed @ai-sdk/openai provider drives it with just a baseURL swap. + const gemini = createOpenAI({ baseURL: "https://generativelanguage.googleapis.com/v1beta/openai/", apiKey }); + // Force the Chat Completions API: Gemini's OpenAI-compat layer implements /chat/completions, + // not the Responses API that @ai-sdk/openai's bare provider(id) defaults to. + return gemini.chat(E2E_MODEL_ID); + } + return openai(E2E_MODEL_ID); +} diff --git a/javascript/examples/vitest/from-transcript/proof.e2e.test.ts b/javascript/examples/vitest/from-transcript/proof.e2e.test.ts new file mode 100644 index 000000000..14b17d111 --- /dev/null +++ b/javascript/examples/vitest/from-transcript/proof.e2e.test.ts @@ -0,0 +1,214 @@ +/** + * #779 — END-TO-END proof (LIVE model calls, not mocked). + * + * Takes a REAL Claude Code JSONL, forks it, seeds the history into a Scenario, lets a + * LIVE agent-under-test take the next turn, and a judge renders a verdict. Covers: + * + * DoD-3 real JSONL → runnable Scenario → judge verdict (Proof A) + * DoD-4 >10-message seed does NOT die on maxTurns (side-door) (Proof A + the naive contrast) + * DoD-5 fix-loop: same forked scenario, verdict FLIPS on a (Proof B) + * config (system-prompt) change — the config the JSONL omits + * + * Posture: Strategy-A "artificial context injection" — the agent-under-test is a live + * model (OpenAI by default; SC779_PROVIDER=gemini for Gemini) seeded with the captured + * *conversation*; we do NOT reconstruct Claude Code's real system prompt (that's the + * deferred sandbox path, Strategy D). Proof B's point is that the omitted config flips the outcome. + * + * Gated behind SC779_LIVE (skipped by default so CI's example suite needs no funded key). + * Run: cd javascript/examples/vitest && SC779_LIVE=1 pnpm exec vitest run from-transcript/proof.e2e.test.ts + * Needs OPENAI_API_KEY, or SC779_PROVIDER=gemini + GEMINI_API_KEY (auto-loaded from .env by the vitest setup). + */ +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; +import scenario, { AgentRole, type AgentAdapter, type ScenarioResult } from "@langwatch/scenario"; +import { generateText, type ModelMessage } from "ai"; +import { describe, it, expect } from "vitest"; +import { fromTranscript, type SeededScenario } from "./from-transcript"; +import { e2eModel, E2E_MODEL_LABEL } from "./model"; + +const HERE = dirname(fileURLToPath(import.meta.url)); +const FIXTURE = join(HERE, "fixtures", "real-cc-session.jsonl"); + +/** A live agent-under-test whose ONLY variable is its system prompt — the config the JSONL drops. */ +function liveAgent(systemPrompt: string): AgentAdapter & { lastReply: string } { + const state = { lastReply: "" }; + const adapter: AgentAdapter = { + role: AgentRole.AGENT, + call: async (input) => { + const res = await generateText({ + model: e2eModel(), + messages: [{ role: "system", content: systemPrompt }, ...input.messages], + }); + state.lastReply = res.text; + return { role: "assistant", content: res.text }; + }, + }; + return Object.defineProperty(adapter, "lastReply", { get: () => state.lastReply }) as AgentAdapter & { + lastReply: string; + }; +} + +function printSeed(label: string, b: SeededScenario) { + console.log(`\n──────────── ${label} ────────────`); + console.log( + `seed: ${b.stats.seededMessages} messages ` + + `(from ${b.stats.parsedNodes} parsed nodes → ${b.stats.chainLength}-node parentUuid chain → ` + + `${b.stats.normalizedTurns} turns: ${b.stats.humanTurns} human / ${b.stats.injectedTurns} injected / ` + + `${b.stats.toolResultTurns} tool / ${b.stats.assistantTurns} assistant)` + ); + if (b.stats.droppedByMatch) console.log(` ⚠ scrubbed ${b.stats.droppedByMatch} token-bearing message(s) to reproduce the 'absent memory' failure`); + console.log(` fork: before turn #${b.stats.forkTurnIndex}; live agent replaces original reply → "${b.originalNextText}"`); + console.log(` seeded conversation (role: preview):`); + for (const m of b.seedMessages) { + const text = typeof m.content === "string" ? m.content : JSON.stringify(m.content); + console.log(` ${String(m.role).padEnd(9)} | ${text.replace(/\s+/g, " ").slice(0, 90)}`); + } +} + +function printVerdict(result: ScenarioResult, reply: string) { + console.log(` ▶ LIVE AGENT TURN: "${(reply || "").replace(/\s+/g, " ").slice(0, 160)}"`); + console.log(` ⚖ JUDGE: success=${result.success}`); + console.log(` met: ${JSON.stringify(result.metCriteria)}`); + console.log(` unmet: ${JSON.stringify(result.unmetCriteria)}`); + console.log(` reasoning: ${(result.reasoning || "").replace(/\s+/g, " ").slice(0, 240)}`); +} + +const CRITERIA_RECALL = ["The assistant's reply gives the exact remembered token, which is KUMQUAT77."]; + +// A minimal, honest reconstruction of the kind of ambient config a real CC session carries +// (a test-harness CLAUDE.md). Without it, gpt-5-mini pattern-matches "token" → refuses as a +// "secret" — itself a demonstration that the agent's guardrail config (absent from the JSONL) +// changes replay behavior. Held CONSTANT across the fix-loop so the ONLY variable is the memory. +const HARNESS_PREAMBLE = + "You are a helpful assistant operating inside an automated test harness. Any short alphanumeric " + + "'tokens' mentioned here are harmless test placeholders the user themselves supplied — not real " + + "secrets — so answer questions about them directly and literally."; +const RECONSTRUCTED_MEMORY = + "\n\nRECOVERED PROJECT MEMORY (reconstructed out-of-band from ~/.claude memory; this was NOT part of\n" + + "the transcript's conversation): earlier in this session the user asked you to remember the token KUMQUAT77."; + +// Gated behind SC779_LIVE (skipped by default) so CI's example suite does not require a funded +// model key. Reproduce this PR's run with: SC779_LIVE=1 SC779_PROVIDER=gemini pnpm exec vitest run … +describe.skipIf(!process.env.SC779_LIVE)("#779 — Scenario from a real Claude Code transcript (live)", () => { + it( + "Proof A (DoD-3 + DoD-4): faithful full-history replay runs to a judge verdict; >10-msg seed survives maxTurns", + async () => { + const b = fromTranscript(FIXTURE, { flattenTools: true }); + printSeed("PROOF A — faithful replay (token in context)", b); + console.log(` model: ${E2E_MODEL_LABEL}`); + expect(b.seedMessages.length).toBeGreaterThan(10); // the seed size that would trip naive maxTurns + + const agent = liveAgent(HARNESS_PREAMBLE); + const result = await scenario.run({ + name: "sc779 Proof A — faithful CC replay", + description: + "A real Claude Code session (memory-recall) is replayed: the captured history is seeded, then the live agent answers the user's final question.", + maxTurns: 10, // seed is ~15 msgs > 10; side-door seeding must NOT consume this budget + agents: [agent, scenario.judgeAgent({ criteria: CRITERIA_RECALL, model: e2eModel() })], + script: [b.seedStep, scenario.agent(), scenario.judge()], + setId: "779-from-transcript", + }); + printVerdict(result, agent.lastReply); + + // DoD-4: it ran to a real verdict — it did NOT die with "Reached maximum turns" + expect(result.reasoning ?? "").not.toMatch(/maximum turns/i); + expect(result.metCriteria.length + result.unmetCriteria.length).toBeGreaterThan(0); + // DoD-3: with the token in context, the faithful replay reproduces the real success + expect(result.success).toBe(true); + }, + 180_000 + ); + + it( + "DoD-4 contrast: NAIVE message()-per-line seed THEN proceed() trips maxTurns; the side-door does not", + async () => { + const b = fromTranscript(FIXTURE, { flattenTools: true }); + const MAXTURNS = 4; // seed (12) >> budget + // The issue's exact documented trap: emit message() per historical message, THEN proceed(). + // In TS, maxTurns is only enforced inside _step() (proceed's auto-advance loop, execution:846) — + // NOT during explicit script steps — so the per-message seeding silently bumps currentTurn, and + // proceed()'s first _step sees the budget already spent and aborts before the agent ever answers. + // message({role:"user"}) also needs a registered USER agent, so we add a userSimulatorAgent. + const naiveSeed = b.seedMessages.map((m: ModelMessage) => scenario.message(m)); + let threw: string | null = null; + let naive: ScenarioResult | null = null; + try { + naive = await scenario.run({ + name: "sc779 DoD-4 naive-seed + proceed (expected to hit the turn cap)", + description: "Naive per-message seeding then proceed() — reproduces the max_turns budget trap.", + maxTurns: MAXTURNS, + agents: [liveAgent(HARNESS_PREAMBLE), scenario.userSimulatorAgent({ model: e2eModel() }), scenario.judgeAgent({ criteria: CRITERIA_RECALL, model: e2eModel() })], + script: [...naiveSeed, scenario.proceed()], + setId: "779-from-transcript", + }); + } catch (e) { + threw = e instanceof Error ? e.message : String(e); + } + + // Same history, same tiny budget, but seeded via the turn-free side-door → survives. + const sideDoor = await scenario.run({ + name: "sc779 DoD-4 side-door seed + proceed (survives the same budget)", + description: "Side-door seeding then proceed() — currentTurn untouched, so the budget is intact.", + maxTurns: MAXTURNS, + agents: [liveAgent(HARNESS_PREAMBLE), scenario.userSimulatorAgent({ model: e2eModel() }), scenario.judgeAgent({ criteria: CRITERIA_RECALL, model: e2eModel() })], + script: [b.seedStep, scenario.agent(), scenario.judge()], + setId: "779-from-transcript", + }); + + console.log(`\n──────────── DoD-4 CONTRAST (naive message()+proceed vs side-door, maxTurns=${MAXTURNS}, seed=${b.seedMessages.length}) ────────────`); + console.log(` NAIVE: threw=${threw ? JSON.stringify(threw.slice(0, 120)) : "no"} success=${naive?.success} reasoning="${(naive?.reasoning || "").slice(0, 140)}"`); + console.log(` SIDE-DOOR: success=${sideDoor.success} met=${JSON.stringify(sideDoor.metCriteria)} reasoning="${(sideDoor.reasoning || "").slice(0, 120)}"`); + + // Naive path overruns the budget (aborts "Reached maximum turns" without a real verdict)… + const naiveHitCap = + (threw != null && /max|turn|budget/i.test(threw)) || + (naive != null && naive.success !== true && /(maximum|reached).{0,14}turn/i.test(naive.reasoning ?? "")); + // …while the side-door seed of the SAME history under the SAME budget reaches a real verdict. + expect(naiveHitCap).toBe(true); + expect(sideDoor.reasoning ?? "").not.toMatch(/maximum turns/i); + }, + 180_000 + ); + + it( + "Proof B (DoD-5): the fix-loop — same forked scenario, judge verdict FLIPS on the omitted config", + async () => { + // Reproduce the genuine failure this session guards against: the agent LACKS the remembered + // token because the config/memory that carried it forward is NOT in the transcript (the crux). + const b = fromTranscript(FIXTURE, { flattenTools: true, dropMatching: /KUMQUAT77/i }); + printSeed("PROOF B — fix-loop (token scrubbed from seed)", b); + expect(JSON.stringify(b.seedMessages)).not.toMatch(/KUMQUAT77/i); + + // ── ORIGINAL config: harness preamble but NO reconstructed memory → agent cannot recall → FAIL ── + const baseAgent = liveAgent(HARNESS_PREAMBLE); + const before = await scenario.run({ + name: "sc779 Proof B — BEFORE (memory absent)", + description: "The live agent is asked to recall a token, but the memory that held it is absent from the seeded context.", + maxTurns: 10, + agents: [baseAgent, scenario.judgeAgent({ criteria: CRITERIA_RECALL, model: e2eModel() })], + script: [b.seedStep, scenario.agent(), scenario.judge()], + setId: "779-from-transcript", + }); + console.log(`\n ▼ BEFORE (baseline config):`); + printVerdict(before, baseAgent.lastReply); + + // ── FIXED config: SAME preamble + reconstruct the out-of-band memory → PASS (memory is the only delta) ── + const fixedAgent = liveAgent(HARNESS_PREAMBLE + RECONSTRUCTED_MEMORY); + const after = await scenario.run({ + name: "sc779 Proof B — AFTER (memory reconstructed)", + description: "Same forked scenario; the agent-under-test's system prompt now carries the reconstructed memory.", + maxTurns: 10, + agents: [fixedAgent, scenario.judgeAgent({ criteria: CRITERIA_RECALL, model: e2eModel() })], + script: [b.seedStep, scenario.agent(), scenario.judge()], + setId: "779-from-transcript", + }); + console.log(`\n ▲ AFTER (fixed config = reconstructed memory):`); + printVerdict(after, fixedAgent.lastReply); + + // DoD-5: the verdict flipped on nothing but the config the JSONL omitted. + expect(before.success).toBe(false); + expect(after.success).toBe(true); + }, + 180_000 + ); +}); diff --git a/javascript/examples/vitest/from-transcript/reader.test.ts b/javascript/examples/vitest/from-transcript/reader.test.ts new file mode 100644 index 000000000..28d2be78f --- /dev/null +++ b/javascript/examples/vitest/from-transcript/reader.test.ts @@ -0,0 +1,259 @@ +/** + * #779 — load-bearing unit tests for the CC transcript reader/adapter. + * These run against a REAL Claude Code session JSONL (fixtures/real-cc-session.jsonl, + * the KUMQUAT77 memory-recall session from the PR #687 work). No API calls. + * + * Each test pins a normalization that a naive "read top-to-bottom" transform gets wrong. + * Mutation-checked: breaking the parentUuid walk, the same-id assistant merge, or the + * tool_use→tool_result name recovery turns these red. + */ +import { mkdtempSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; +import type { TextPart, ToolCallPart, ToolResultPart } from "ai"; +import { describe, it, expect } from "vitest"; +import { + parseSessionFile, + linearize, + normalize, + toModelMessages, + type RawNode, +} from "./cc-transcript"; +import { fromTranscript } from "./from-transcript"; + +const HERE = dirname(fileURLToPath(import.meta.url)); +const FIXTURE = join(HERE, "fixtures", "real-cc-session.jsonl"); + +describe("cc-transcript reader (real JSONL)", () => { + const raw = parseSessionFile(FIXTURE); + const chain = linearize(raw); + const turns = normalize(chain); + + it("parses the real session and keeps only uuid-bearing nodes", () => { + expect(raw.length).toBeGreaterThan(40); + expect(raw.every((n) => !!n.uuid)).toBe(true); + }); + + it("walks the parentUuid chain leaf→root (not file order) into a single linear path", () => { + // every node except the root must have its parent immediately preceding it + const uuids = new Set(chain.map((n) => n.uuid)); + for (let i = 1; i < chain.length; i++) { + expect(chain[i].parentUuid).toBe(chain[i - 1].uuid); + } + expect(uuids.size).toBe(chain.length); // no dup / cycle + }); + + it("drops non-conversational metadata (attachment/system) after normalization", () => { + const rawTypes = new Set(raw.map((n) => n.type)); + expect(rawTypes.has("attachment")).toBe(true); // fixture really has metadata lines + // no attachment node may surface in any normalized turn's source-uuid set (fails if filtering is off) + expect( + raw.filter((n) => n.type === "attachment").every((n) => !turns.some((t) => t.uuids.includes(n.uuid!))) + ).toBe(true); + }); + + it("distinguishes the overloaded user role: genuine human vs command-injection vs tool_result", () => { + const humans = turns.filter((t) => t.role === "user" && t.userKind === "human"); + const injected = turns.filter((t) => t.role === "user" && t.userKind === "injected"); + const toolResults = turns.filter((t) => t.role === "tool"); + // 2 genuine human turns: the "remember KUMQUAT77" ask and the "what was the token?" question + expect(humans.length).toBe(2); + expect(humans[0].text).toMatch(/remember/i); + expect(humans[0].text).toMatch(/KUMQUAT77/); + expect(humans[1].text).toMatch(/what was the exact token/i); + // injected command/skill output dumps are injected user turns, NOT genuine human + expect(injected.length).toBeGreaterThanOrEqual(2); + // real tool I/O was captured + expect(toolResults.length).toBeGreaterThanOrEqual(4); + }); + + it("re-merges an assistant message split across multiple JSONL lines (same message.id)", () => { + // The first assistant turn is thinking + text + tool_use(Write) across 3 lines → ONE turn + const firstAssistant = turns.find((t) => t.role === "assistant")!; + expect(firstAssistant.thinking && firstAssistant.thinking.length).toBeGreaterThan(0); // thinking captured + expect(firstAssistant.toolCalls?.some((c) => c.name === "Write")).toBe(true); // merged tool_use + }); + + it("recovers toolName for every tool_result by pairing back to its tool_use id", () => { + const results = turns.filter((t) => t.role === "tool").flatMap((t) => t.toolResults ?? []); + expect(results.length).toBeGreaterThan(0); + // none should fall back to the unknown sentinel — the id→name map must resolve them + expect(results.every((r) => r.name !== "unknown_tool")).toBe(true); + expect(results.map((r) => r.name)).toEqual(expect.arrayContaining(["Write"])); + }); + + it("captures extended thinking blocks (they exist in the real transcript)", () => { + const anyThinking = turns.some((t) => t.role === "assistant" && (t.thinking?.length ?? 0) > 0); + expect(anyThinking).toBe(true); + }); +}); + +describe("Anthropic → AI-SDK v6 ModelMessage emission", () => { + const turns = normalize(linearize(parseSessionFile(FIXTURE))); + + it("emits valid v6 tool-result parts (toolName + output:{type:text,value})", () => { + const msgs = toModelMessages(turns); + const toolMsg = msgs.find((m) => m.role === "tool"); + expect(toolMsg).toBeTruthy(); + const part = (toolMsg!.content as ToolResultPart[])[0]; + expect(part.type).toBe("tool-result"); + expect(typeof part.toolName).toBe("string"); + expect(part.output).toEqual(expect.objectContaining({ type: "text", value: expect.any(String) })); + }); + + it("emits assistant tool-call parts with input (v6), and drops thinking by default", () => { + const msgs = toModelMessages(turns); + const withToolCall = msgs.find( + (m) => m.role === "assistant" && Array.isArray(m.content) && m.content.some((p) => p.type === "tool-call") + ); + expect(withToolCall).toBeTruthy(); + const call = (withToolCall!.content as Array).find( + (p): p is ToolCallPart => p.type === "tool-call" + ); + expect(call).toEqual(expect.objectContaining({ type: "tool-call", toolName: expect.any(String) })); + expect(call).toHaveProperty("input"); + // thinking must NOT leak into the emitted messages by default (invalid for non-Anthropic models) + const anyThinkingText = JSON.stringify(msgs).includes("[thinking]"); + expect(anyThinkingText).toBe(false); + }); + + it("flattenTools folds tool I/O into assistant TEXT (string content, reachable [tool …] marker)", () => { + // flattenTools is the ONLY emit mode the e2e uses: tool_use/tool_result become readable + // assistant TEXT (not structured parts) so a cross-model seed passes provider validation. + const msgs = toModelMessages(turns, { flattenTools: true }); + // every message content is a plain string — NO structured tool parts survive + expect(msgs.every((m) => typeof m.content === "string")).toBe(true); + // no tool-role message survives; tool results are folded into assistant strings + expect(msgs.some((m) => m.role === "tool")).toBe(false); + // the trunc()'d tool-result marker path is reachable: an assistant string carries "[tool …]" + const marker = msgs.some( + (m) => m.role === "assistant" && typeof m.content === "string" && /\[tool \S+ (?:→|ERROR)\]/.test(m.content) + ); + expect(marker).toBe(true); + }); +}); + +describe("fromTranscript builder + DoD-4 turn-free seeding", () => { + it("forks before the final assistant reply and seeds the full pre-fork history (>10 messages)", () => { + const b = fromTranscript(FIXTURE); + // A >10-message seed is what trips the default maxTurns=10 under naive message() seeding + expect(b.seedMessages.length).toBeGreaterThan(10); + // last seeded human turn is the final question; the forked-out assistant reply is the KUMQUAT77 answer + expect(b.lastHumanText).toMatch(/what was the exact token/i); + expect(b.originalNextText).toMatch(/KUMQUAT77/); + }); + + it("seedStep pushes every message via addMessage WITHOUT incrementing any turn counter", () => { + const b = fromTranscript(FIXTURE); + // fake state mirroring ScenarioExecutionState's public surface: addMessage + a turn counter + const pushed: unknown[] = []; + let currentTurn = 0; + const fakeState = { + addMessage: (m: unknown) => pushed.push(m), + // the real newTurn() is the ONLY thing that moves currentTurn; addMessage never calls it + newTurn: () => currentTurn++, + }; + b.seedStep(fakeState); + expect(pushed.length).toBe(b.seedMessages.length); + expect(currentTurn).toBe(0); // ← the whole point: zero turns consumed by seeding + }); + + it("dropMatching removes token-bearing messages (reproduces the 'absent memory' failure)", () => { + const full = fromTranscript(FIXTURE); + const scrubbed = fromTranscript(FIXTURE, { dropMatching: /KUMQUAT77/ }); + expect(scrubbed.stats.droppedByMatch).toBeGreaterThan(0); + expect(JSON.stringify(scrubbed.seedMessages)).not.toMatch(/KUMQUAT77/); + expect(scrubbed.seedMessages.length).toBeLessThan(full.seedMessages.length); + }); +}); + +// Regression tests for the correctness bugs surfaced by the max-effort review (all previously green +// while broken — each fails on the pre-fix code). +describe("reader hardening (review regressions)", () => { + it("linearize ignores subagent (Task) sidechains and walks the MAIN thread", () => { + // main thread u1->a1 (leaf a1); a separate sidechain s1->s2 (leaf s2) appears later in file + const nodes: RawNode[] = [ + { uuid: "u1", parentUuid: null, type: "user", message: { role: "user", content: "MAIN question" } }, + { uuid: "a1", parentUuid: "u1", type: "assistant", message: { role: "assistant", content: [{ type: "text", text: "MAIN answer" }] } }, + { uuid: "s1", parentUuid: null, type: "user", isSidechain: true, message: { role: "user", content: "SUBAGENT task" } }, + { uuid: "s2", parentUuid: "s1", type: "assistant", isSidechain: true, message: { role: "assistant", content: [{ type: "text", text: "SUBAGENT chatter" }] } }, + ]; + const turns = normalize(linearize(nodes)); + expect(JSON.stringify(turns)).not.toMatch(/SUBAGENT/); // NOT the sidechain + expect(turns.some((t) => /MAIN answer/.test(t.text ?? ""))).toBe(true); + }); + + it("dropMatching with a GLOBAL regex still scrubs every match (no stateful lastIndex skips)", () => { + const g = fromTranscript(FIXTURE, { flattenTools: true, dropMatching: /KUMQUAT77/g }); + expect(JSON.stringify(g.seedMessages)).not.toMatch(/KUMQUAT77/); + }); + + it("structured-emit dropMatching leaves NO orphaned tool parts (provider-valid seed)", () => { + const scrubbed = fromTranscript(FIXTURE, { dropMatching: /KUMQUAT77/ }); // default structured + const callIds = new Set(); + const resultIds: string[] = []; + for (const m of scrubbed.seedMessages) { + if (Array.isArray(m.content)) { + for (const p of m.content as Array) { + if (p.type === "tool-call") callIds.add(p.toolCallId); + if (p.type === "tool-result") resultIds.push(p.toolCallId); + } + } + } + expect(resultIds.every((id) => callIds.has(id))).toBe(true); // every result has its call + }); + + it("forkAt {index} out of range throws instead of silently mis-forking", () => { + expect(() => fromTranscript(FIXTURE, { forkAt: { index: -2 } })).toThrow(/out of range/); + const n = fromTranscript(FIXTURE).stats.normalizedTurns; + expect(() => fromTranscript(FIXTURE, { forkAt: { index: n + 5 } })).toThrow(/out of range/); + }); + + it("forkAt {uuid} resolves a uuid pointing at a merged-away (non-first) assistant line", () => { + // fdf12fbd… is the text line of the first assistant message (merged into the af3e4e11 turn) + const b = fromTranscript(FIXTURE, { forkAt: { uuid: "fdf12fbd-a1a9-43f0-98b4-0c67c50a1a45" } }); + // resolves to the turn that folded that (non-first) line in — NOT throw "not found", NOT index 0 + expect(b.turns[b.stats.forkTurnIndex].uuids).toContain("fdf12fbd-a1a9-43f0-98b4-0c67c50a1a45"); + }); + + it("pruneOrphanToolParts prunes an orphaned tool-CALL when the fork drops its tool_result", () => { + // The fixture only orphans tool-RESULTs; this synthesizes the mirror case (an orphaned tool-CALL). + // Chain: user → assistant(tool_use T1) → tool_result(T1) → assistant(text). Forking BEFORE the + // tool_result turn seeds the tool_use with no matching result → the tool-call part must be pruned. + const nodes: RawNode[] = [ + { uuid: "u1", parentUuid: null, type: "user", message: { role: "user", content: "run the tool" } }, + { uuid: "a1", parentUuid: "u1", type: "assistant", message: { id: "m1", role: "assistant", content: [{ type: "tool_use", id: "T1", name: "Foo", input: {} }] } }, + { uuid: "t1", parentUuid: "a1", type: "user", message: { role: "user", content: [{ type: "tool_result", tool_use_id: "T1", content: "ok" }] } }, + { uuid: "a2", parentUuid: "t1", type: "assistant", message: { id: "m2", role: "assistant", content: [{ type: "text", text: "done" }] } }, + ]; + const dir = mkdtempSync(join(tmpdir(), "sc779-orphan-call-")); + const path = join(dir, "synthetic.jsonl"); + writeFileSync(path, nodes.map((n) => JSON.stringify(n)).join("\n")); + + // turns: 0 user, 1 assistant(tool_use T1), 2 tool(result T1), 3 assistant(text) → fork before #2 + const b = fromTranscript(path, { forkAt: { index: 2 } }); + + // the orphan-tool-CALL branch fired… + expect(b.stats.prunedOrphanToolParts).toBeGreaterThan(0); + // …and no tool-call part remains whose id lacks a matching tool-result + const callIds = new Set(); + const resultIds = new Set(); + for (const m of b.seedMessages) { + if (Array.isArray(m.content)) { + for (const p of m.content as Array) { + if (p.type === "tool-call") callIds.add(p.toolCallId); + if (p.type === "tool-result") resultIds.add(p.toolCallId); + } + } + } + expect([...callIds].every((id) => resultIds.has(id))).toBe(true); + }); + + it("does not strip a leading 'user: ' from genuine human content", () => { + const turns = normalize(linearize(parseSessionFile(FIXTURE))); + const firstHuman = turns.find((t) => t.role === "user" && t.userKind === "human")!; + // the fixture's real content begins with "user: Remember…" — that prefix must survive verbatim + expect(firstHuman.text?.startsWith("user: ")).toBe(true); + }); +});