feat(dispatch): add Sonnet 5.5 and GPT-6.1 Sol pins - #334
Conversation
PR SummaryMedium Risk Overview For Cursor, native probes showed Sonnet 5.5 bracket selectors fall back while exact model IDs resolve correctly. The pin-mapping decision and dispatch docs now allow explicit exact-ID registry entries with per-mapping probe evidence—catalog visibility alone cannot pass through arbitrary IDs. Probe fixtures, runbook references, and skill/provider guidance are updated accordingly. Codex gains new Reviewed by Cursor Bugbot for commit 6c18f4f. Bugbot is set up for automated code reviews on this repo. Configure here. |
Resolve skills.test.ts pin conflicts (main's plan-writing 1.2.34 plus wave 3 bumps) and move oat-project-implement to 2.3.16 and oat-dispatch-subagents to 1.2.12 above main's 2.3.15/1.2.11.
Follows the version files #334 (98d1d52) changed: the five public package.json files and packages/cli/assets/public-package-versions.json. The sync manifest's oatVersion is a sync stamp, not a lockstep version, and is left unchanged. Verified: release:check-versions exit 0 against origin/main 98d1d52; release:validate exit 0 (5 packages at 0.3.11); scoped release and bundle tests pass.
… validate-only dispatch record (wave 3, lockstep 0.3.11) (#336) * chore(pjm): set wave 3 lead items Raise BL-260927-expose-a-scoped-template to high as the wave lead and give BL-260718-support-fumadocs-in-oat-docs acceptance criteria: nav sync writes Fumadocs meta.json from authored index.md Contents maps. * chore(pjm): file recon issue #333 as three backlog items The packet-validator contract fix joins wave 3 at high priority; Codex capacity recovery and controller setup friction are tracked at medium and low. * chore(pjm): scope recon recovery to a note and one retry Mixed native and CLI continuation records move to a deferred item until the friction recurs. * chore(pjm): file two lifecycle routing items from wave 3 recon * chore(oat): scaffold backlog-wave-3 * chore(oat): enable all-phase gates for backlog-wave-3 * chore(oat): persist autonomous explainer intent Also records the managed/high dispatch policy chosen for this wave. * chore(oat): update autonomous execution learnings * chore(oat): capture quick-start discovery for backlog-wave-3 * chore(oat): complete quick-start discovery for backlog-wave-3 * chore(oat): draft backlog-wave-3 plan for review * chore(oat): apply plan artifact review findings * chore(oat): update autonomous execution learnings * chore(oat): apply plan re-review findings * chore(oat): record plan artifact review disposition * chore(oat): record plan review artifact * chore(oat): record gate review in project log * chore(oat): resolve plan gate and complexity review findings * chore(oat): record plan review artifact * chore(oat): record gate review in project log * chore(oat): record plan gate attempt 2 and stop at the QS-12 boundary * chore(oat): update quick-start artifacts for backlog-wave-3 * chore(oat): initialize backlog-wave-3 implementation * chore(oat): add the default phase recovery ledger * feat(p01-t01): share one template resolver in repository, user, bundle order Move the PJM template resolver to commands/shared/template-source.ts as resolveTemplate (repository, user, bundle) and drop "PJM" from its messages. pjm init, backlog new and decision new call it unchanged. Project scaffold and lite promotion now call it with templatesRoot=<repo>/.oat/templates instead of the user-first resolveTemplateSource, which is deleted. The bundle root may be passed lazily so scaffold still resolves assets only when both tiers miss. Unchanged on purpose: cleanup/project/project.ts (repository or inline) and project/log/append.ts (bundle only). Neither copies a lifecycle template that a user-scope install lacks. Failing first: flipped scaffold.test.ts "uses a repo template before a differing user template" failed before the change with "expected '# repo-first USER-TEMPLATE...' to contain 'REPO-TEMPLATE'". The partial-tier scaffold test also encoded user-first order; it now proves repo plan, user discovery and bundle state in one scaffold. Verify: HOME=$(mktemp -d) vitest run src/commands/{shared,project,pjm, backlog,decision}: 126 files, 2365 tests passed; pnpm type-check exit 0. * feat(p01-t02): add oat template resolve Add `oat template resolve <name> [--output <path>]` (global --json). It resolves through the shared repository, user, bundle resolver and reports name, found, tier and path. path is null for the bundle tier, so no package-manager path leaks. <name> takes plan or plan.md and rejects separators and "..". --output copies the content, replacing an existing file, and creates no parent directories. A miss exits 1 and names the three tiers. Outside a git repository the repository tier is skipped. The shared resolver gains TemplateNotFoundError (same message) so the command can tell a miss from a read failure. Failing first: with the help snapshots and template/resolve.test.ts written before the command, vitest failed 3 snapshot tests (root help, template help, template resolve help) and the resolve suite failed to import the missing module. Verify: build exit 0; HOME=$(mktemp -d) vitest run src/commands/template src/commands/help-snapshots.test.ts: 77 tests passed; oat-docs check exit 0; type-check and lint exit 0. Branch-CLI probe (issue #296): temp git repo with no .oat/templates and HOME holding ~/.oat/templates/plan.md -> tier user, --output copied it; discovery -> tier bundle, path none; nope -> exit 1 found:false; ../x -> exit 1; --output into a missing dir -> exit 1. * feat(p01-t03): route lifecycle skills through oat template resolve Lifecycle skills copy templates with `oat template resolve <name> --output "$PROJECT_PATH/<file>"` instead of reading .oat/templates/, so a user-scope-only install works. Unconditional copies are straight substitutions; fill-if-missing steps (quick-start, import-plan plan.md) keep their condition in prose. Retro now creates references/ first because --output creates no directories. Skills: retro, design (spec + 3 design copies), spec, discover (state, discovery), plan (copy + overwrite), summary, quick-start (resume, scaffold note, discovery, design, plan), import-plan (plan, implementation), promote-spec-driven, implement plan-and-resume, new, repo-improve, pjm-decision. cursor-cloud-projects: templates resolve repository, user, bundle via the CLI; skills and scripts keep user-first order, and the staleness rationale now covers only skills and scripts. Ideas skills are unchanged. Docs: cursor-cloud, file-locations, tool-packs, troubleshooting. Pins: skills.test.ts quick-start/import-plan rewrite markers and every bumped version; review-skill-contracts.test.ts quick-start resume pins. Bumps (patch): retro 1.0.8, design 2.3.7, spec 2.0.4, discover 2.2.8, plan 1.4.16, summary 1.5.7, quick-start 2.3.17, import-plan 1.4.19, promote-spec-driven 1.2.4, implement 2.3.15, cursor-cloud-projects 1.1.3, new 1.4.2, repo-improve 2.1.6, pjm-decision 1.1.2. Failing first: after the text edits, vitest src/validation failed 16 tests (stale version pins and moved rewrite markers) and init/tools/shared failed the quick-start resume pin; both pass now. Verify: rg "\.oat/templates/" .agents/skills (non-test) leaves only descriptive matches; oat:validate-skills exit 0; HOME=$(mktemp -d) vitest run src/validation src/commands/init/tools/shared: 899 passed; node --test oat-project-implement/tests exit 0; oat-docs check exit 0; oxlint .agents/skills exit 0; oat sync --scope project --dry-run: no changes. * chore(oat): record p01 task ledger before review * chore(oat): receive p01 review and add p01-t04 * fix(p01-t04): close p01 review findings M1: oat-project-retro and oat-project-summary curate oat access one subcommand at a time, so both now grant Bash(oat template:*) for their template copy; retro also grants Bash(mkdir:*) for creating references/. skills.test.ts pins both grants next to the Bash(oat tools:*) pin. The other p01-t03 skills need nothing: discover and implement grant Bash(oat:*), the rest grant bare Bash, and design/spec/plan's git-only list predates p01 (review scoped it out). Both skills were already bumped on this branch (1.0.8, 1.5.7), so no further bump. The summary allowed-tools edit changed that line's autonomy prompt-site key, so .agents/docs/autonomy-contract.md maps the new key a8983fec040c -> NG in place of 35cb2ea1d677 -> NG (same disposition). L1: tool-packs.md moves the oat template resolve sentence above the fallback paragraph, so "Useful options" stays attached to oat pjm init. Failing first: the new pin failed before any grant ("expected ... to contain 'Bash(oat template:*)'"), then failed again on 'Bash(mkdir:*)' with only the oat template grants added; it passes with both. Verify: HOME=$(mktemp -d) vitest run src/validation: 391 passed; oat:validate-skills, check:skill-bumps and oat-docs check exit 0. * chore(oat): bookkeeping after p01 pass * chore(oat): record p01 review artifact * chore(oat): record gate review in project log * chore(oat): receive p01 gate findings and add p01-t05 * fix(p01-t05): close p01 gate findings M1: oat-project-design, oat-project-spec and oat-project-plan now run oat template resolve, so each allowed-tools list grants Bash(oat template:*). The p01-t04 allowlist contract test in skills.test.ts now covers all three alongside retro and summary. The design and plan header edits changed their autonomy prompt-site key, so .agents/docs/autonomy-contract.md maps f9e319c9b223 -> NG in place of 5eb3949f32e1 -> NG (same disposition); spec is not an autonomy root. All three skills already carry their branch bumps (2.3.7, 2.0.4, 1.4.16). L1: file-locations.md, tool-packs.md and troubleshooting.md now limit the repository, user, bundle claim to project lifecycle and PJM templates and note that idea templates follow the ideas scope. The resolver is unchanged. Failing first: the extended test failed before any grant on design ("expected 'name: oat-project-design...' to contain 'Bash(oat template:*)'"). Adding grants one at a time, it then failed on spec, then on plan, and passed once all three were granted. Verify: HOME=$(mktemp -d) vitest run src/validation: 391 passed; oat:validate-skills, check:skill-bumps and oat-docs check exit 0. * chore(oat): bookkeeping after p01 gate pass * feat(p02-t01): write Fumadocs meta.json from index.md Contents maps `oat docs nav sync` detects the framework: mkdocs.yml keeps the MkDocs path; a Fumadocs source.config.* writes one strict meta.json per docs directory with an index.md. Pages follow the Contents order with no "..." rest entry, the root lists "index" first, subfolders appear by name, and cross-folder Contents links become Fumadocs link entries. The folder title comes from the index.md frontmatter title, else its first H1. Existing files are compared by meaning, so a rerun writes nothing even after reformatting. Unlisted pages and folders are reported in human output and --json. Failing first: 8 new tests in sync.test.ts failed before the change (framework undefined, no meta.json, no command reporting); 6 existing passed. Neutralize and restore: forcing changed=true failed the rerun and no-changes tests; naming cross-folder pages instead of linking failed the nested-fixture test; both pass after restore. Verify: vitest run src/commands/docs, 14 files, 179 tests pass. * docs(p02-t02): describe Fumadocs nav sync Help text and the docs pages now describe both frameworks: nav sync writes mkdocs.yml nav for MkDocs and strict, committed meta.json files for Fumadocs, reports unlisted pages, and writes cross-folder Contents links as link entries. Adds a Fumadocs Nav Sync section to the docs index contract. Also updates reference/file-locations.md and reference/index.md, which carried the same MkDocs-only wording, and regenerates the tracked apps/oat-docs/index.md manifest for the changed page descriptions. Failing first: the docs nav sync --help inline snapshot failed against the new description and option text before the snapshot update. Verify: vitest help-snapshots + src/commands/docs (15 files, 241 tests) and pnpm --filter oat-docs check pass. * docs(p02-t03): generate Fumadocs navigation for apps/oat-docs Ran the branch CLI `docs nav sync --target-dir apps/oat-docs`: it wrote 11 meta.json files and reported no unlisted pages, so all 70 pages stay in the sidebar and no Contents map needed changes. workflows/skills links into contributing/ and docs-tooling/ became link entries. Failing first: before this commit apps/oat-docs had no meta.json, so the sidebar used Fumadocs' default file-tree order instead of Contents order. Verify: a second run after `oxfmt --write apps/oat-docs/docs` reported written [] and unlisted []. pnpm build:docs passed uncached (0 of 6 cached); the built sidebar order for the root, workflows/projects and contributing matches each meta.json, and the skills link entries render as /open-agent-toolkit/contributing/skills/ and .../docs-tooling/workflows/. pnpm --filter oat-docs check passes. * docs(p02-t04): teach the docs skills about Fumadocs nav sync oat-docs-bootstrap (SKILL.md Sections B and C, AGENTS.md.template) and oat-docs-authoring's Fumadocs contract now say meta.json is generated by `oat docs nav sync` from Contents, committed, and strict. oat-docs-apply changes meta.json only through nav sync and resolves unlisted pages through approved Contents edits. oat-docs-analyze gains read-only Fumadocs meta.json checks (missing, stale, rest entries, unlisted pages) and no longer treats an app without mkdocs.yml as needing migration. Versions: oat-docs-apply 1.3.1 -> 1.3.2, oat-docs-bootstrap 1.2.1 -> 1.2.2, oat-docs-authoring 1.0.1 -> 1.0.2, oat-docs-analyze 1.5.2 -> 1.6.0. None had a bump on this branch versus origin/main; no version pins exist for them. Failing first: prose-only change with no executable contract; before it, the skills described nav sync as MkDocs-only and meta.json as optional hand-authored files, which contradicts p02-t01. Verify: pnpm oat:validate-skills (65 skills), vitest src/validation (6 files, 391 tests), format:root pass; no tests/ dir exists in the oat-docs-* skills. Provider views are symlinks; branch `sync --scope project --dry-run` shows no changes. * chore(oat): record p02 task ledger before review * chore(oat): receive p02 review and add p02-t05 * fix(p02-t05): close p02 review findings M1: add `oat docs nav sync --check`. It computes the navigation, writes nothing, and exits 1 when a meta.json (or mkdocs.yml) would change or a page or folder is unlisted, naming each (`stale`, `unlisted` in --json). apps/oat-docs prebuild runs it after generate-index, so build:docs fails on stale navigation. Help snapshot and nav sync docs updated. M2: Fumadocs Contents maps accept .mdx targets, listed by slug; MkDocs stays .md-only. M3: commands.md no longer implies the hooks write meta.json. L1: apps/oat-docs/AGENTS.md says meta.json is committed and rewritten only by nav sync. L2: tests for the documentation.tooling fallback. L3: inline Markdown is stripped from H1-derived folder titles. oat-docs-analyze (already bumped on this branch) now points analysis at the --check form. Failing first: 7 new tests failed before the change (.mdx slug, inline title, MkDocs check, four --check cases); the help snapshot failed on the new option. The two L2 tests passed against existing code, so they were proven by neutralizing: dropping the root match failed the mismatched-root test, dropping the tooling fallback failed the matching-root test, and dropping the check write guard failed both writes-nothing tests; all pass after restore. Negative control: an unlisted docs/zz-probe.md made `pnpm --filter oat-docs run prebuild` exit 1 naming it; removed, exit 0. Verify: cli build; vitest src/commands/docs + help-snapshots + release contract (16 files, 277 tests); nav sync --check on apps/oat-docs exit 0; oat-docs check; build:docs uncached (0/6 cached) exit 0. * chore(oat): bookkeeping after p02 pass * chore(oat): record p02 dispatch and deviations * chore(oat): record p02 review artifact * chore(oat): record gate review in project log * chore(oat): receive p02 gate findings and add p02-t06 * fix(p02-t06): close p02 gate findings M1: index handling now follows the effective folder metadata. fumadocs-core 16.10.2 gives a root folder (the docs root or a preserved `root: true`) or a folder with a different `pagesIndex` no implicit landing page, so nav sync lists `index` first there. Reachability counts a landing page only when the loader attaches it or `pages` lists it, never just because the file exists. New fumadocs-loader.test.ts loads the generated meta.json through the docs app's installed fumadocs-core loader and asserts every page not reported as unlisted is in the tree, for an ordinary folder and a `root: true` folder, for both sync and --check, and that --check flags an old root-folder meta.json the loader would truncate. docs-index-contract.md notes the rule. M2: oat-docs-analyze grants only Bash(oat docs nav sync --check:*), a prefix that leaves plain (writing) nav sync ungranted; skills.test.ts pins the grant, the command form, and the absence of broader oat or Bash grants. oat-docs-analyze is not an autonomy-contract root, so no prompt site is re-keyed; it is already bumped (1.6.0) on this branch. L1: the analysis command is `oat docs nav sync --check --target-dir <docs-app-dir>`, with the app directory distinguished from its docs/ root. Failing first: the two root: true loader tests and the grant pin failed before the change (the ordinary-folder test passed, as expected). Proofs: ignoring `root: true` in hasImplicitIndex fails both root tests; adding a Bash(oat docs:*) grant fails the pin; both pass after restore. Verify: cli build; vitest src/commands/docs + src/validation (21 files, 583 tests); nav sync --check on apps/oat-docs exit 0; oat:validate-skills, oat-docs check, CLI type-check and check, format:root exit 0. * chore(oat): bookkeeping after p02 gate pass * fix(p03-t01): share review-brief source binding across recon helpers One module (scripts/lib/review-binding.mjs) now owns the review source projection allowlist, the brief-level source union, and each claim's source subset. create-review-brief, the artifact-shape contract, and both validate-packet call sites use it. A verification brief binds when its sources equal the projected union of its claims' sources and each claim's evidence and source subset match. Briefs stay blind. Failing first: on the new two-source fixture (production createReviewBrief plus reconcileLedger), validatePacket reported 5 REVIEW_BRIEF_MISMATCH errors (plus one MATERIAL_COVERAGE_ASSURANCE_EXCEEDED, fixed in p03-t02). Neutralize and restore (review-binding.mjs, once each): - briefSourcesBind -> return true: the projected-union test fails. - compare raw manifest sources instead of the projection: the allowlist test and the copied-descriptor test fail. Restored; recon suite 333/333 pass. The per-claim subset check is implied by the union check plus evidence binding, so neutralizing it alone cannot fail a test; it stays as the documented per-claim rule. Bumps recon 1.1.5 -> 1.1.6 with its pins. * fix(p03-t02): align recon coverage publication with reconciliation Publication now requires only that every claim named by an accepted material coverage finding is not verified, which is the downgrade reconcile-ledger already forces. It no longer also requires the coverage reviewer's per-statement disposition to be gap, so a reviewer may mark statements covered while reporting a missing question. packet-contract.md states the one coverage rule for acceptance, reconciliation, and publication. Failing first: on the two-source fixture (coverage marks every statement covered and reports a material QUESTION_SCOPE_OMISSION), validatePacket reported MATERIAL_COVERAGE_ASSURANCE_EXCEEDED for the downgraded claim. Negative control (a verified claim named by a material finding still fails): integrity-contracts "material coverage gaps prevent verified assurance". Neutralized the guard (if (false && ...)): that test failed (333/334 pass). Restored; recon suite 334/334 pass. * feat(p03-t03): scope recon unresolved issues to claims unresolvedIssues entries are now a closed union: a string (legacy, read as global), { text, claimIds } naming a non-empty list of unique claims the review covers, or { text, scope: 'global' }. Anything else is rejected at artifact acceptance (INVALID_UNRESOLVED_ISSUE) and is never read as no issue. reconcile-ledger and validate-packet share unresolvedIssuesBlockClaim: a scoped issue keeps only its claims below verified; a global issue keeps every covered claim below verified. The conditional-routing predicates (any issue present) are unchanged. packet-contract.md and the worker-contract examples describe the forms. Failing first: the end-to-end test (production create-review-brief, reconcileLedger, validatePacket, renderPacket on the two-source fixture) first threw INVALID_UNRESOLVED_ISSUE while building; with only the shape accepted, it failed with REVIEW_DISPOSITION_MISMATCH alone (x2, the affirmed claims). Flipped the two object-rejection tests; added malformed cases (empty claimIds, non-string ID, uncovered ID, neither, both). Neutralize and restore (once each; restored suite 338/338 pass): - validateUnresolvedIssues -> return: both malformed-scope tests fail. - reconciler drops issueBlocked: scoped-affirmed and global tests fail. - validator back to any-issue-blocks: end-to-end and scoped tests fail. * test(p03-t04): prove recon negative controls against helper output packet-validation's makePacket now builds its verify, adversary, coverage, and redundant-verify briefs with the production createReviewBrief, so the per-code negative tests start from helper output. The excerpt controls use standard packets (which carry those briefs). New controls: - an edited brief statement, evidence excerpt, locator, or source descriptor each fails with REVIEW_BRIEF_MISMATCH and nothing else; - an edited second-source descriptor in the two-source brief fails; - a global semantic issue (object or legacy string) left on a verified claim fails with REVIEW_DISPOSITION_MISMATCH. The material-coverage control already starts from helper briefs (createPacketFixture uses createReviewBrief). Neutralize and restore (once each; restored suite 343/343 pass): - reviewBriefBindsClaim -> return true: 8 tests fail (4 brief edits, two-source descriptor, union, copied descriptor, claim-bearing briefs). - excerpt guard (if (false && ...)): wrong-excerpt and paraphrase fail. - material coverage guard: "material coverage gaps prevent verified assurance" fails. - validator issue guard: the global-issue control fails. * docs(p03-t05): document Codex agent-limit recovery for recon lanes recon: a launch rejected before any child is accepted is a provider/dispatch failure, never a worker failure, and every accepted artifact is kept. retryLimit now means pre-acceptance admission retries per lane: with retryLimit >= 1 the controller makes at most one, after checking completed agents are eligible to be unloaded; 0 allows none. An alternate route is used only when already approved; otherwise the run stops partial (PASS_OMITTED lane gap) and asks for a continuation amendment. No retry or route changes model, effort, role behavior, data authority, output limits, or reviewer blindness; fresh review lanes stay fresh; an accepted lane is never rerun to free capacity. Step 5 points to the dispatch dependency for provider mechanics. The rendered packet labels the limit "pre-acceptance admission retries per lane"; preview attempt math is unchanged. oat-dispatch-subagents provider-codex.md gains the Codex v2 residency note (eligibility, interrupt_agent vs close_agent, queue-only messages pinning completed agents, no archive or delete), citing openai/codex rust-v0.159.2 sources from #333. Appended at the end so the keyed line in autonomy-contract.md does not move. skill-contract pins the guidance and that recon does not copy provider mechanics; replacing the provider-codex pointer fails it (restored). Bumps oat-dispatch-subagents 1.2.10 -> 1.2.11 with its two pins. * chore(oat): record p03 task ledger before review * chore(oat): receive p03 review, record the main merge, add p03-t06 * fix(p03-t06): close p03 review findings H1: validateReviewBindings now binds a verification brief's claim set in both directions. Every disposition already had to bind a brief claim; now every brief claim must also be a distinct ledger claim whose exact projection binds (statement, evidence, its source subset, and the brief-level source union). An injected brief claim fails closed with REVIEW_BRIEF_MISMATCH whether it cites an existing or an injected source. A real ledger claim the reviewer left without a disposition still binds (reconciliation keeps it unresolved), so the existing omitted-claim honest-partial test still passes. Strict disposition equality would have made that packet unpublishable, so it was not used. Briefs stay blind. Failing first: both probe shapes from the review (extra claim on source-1; extra claim plus injected source-3 in the manifest and brief) validated valid=true. Neutralizing the new check: both probe tests fail. M2: render-packet adds a Review Downgrades section listing every claim a review kept below verified (scoped or global issue, coverage finding, non-affirming disposition) with the review's text, key claim or not. The fixture's all-key workaround is reverted (key claims: alpha, gamma). Failing first (renderer at HEAD): 4 tests fail, including a complete packet with a scoped issue on a non-key claim. Neutralizing reviewDowngradeLines: 3 tests fail. L1: the approval preview labels the limit "Pre-acceptance admission retries per lane (at most one used)" and computes worst-case attempts as lanes x (min(retryLimit, 1) + 1). packet-contract.md and the recon docs page state the rule. Failing first (routing at HEAD): the new preview test fails. Restored suite: 349/349 pass. No version bumps. * chore(oat): record p03-t06 before re-review * chore(oat): update autonomous execution learnings * chore(oat): receive p03 round-2 review and add p03-t07 * fix(p03-t07): close p03 round-2 review findings H1: the p03-t06 brief binding now covers every brief mode. Each entry (verification/coverage claims, adversarial provisionalStatements) must be a distinct ledger claim whose exact projection binds; a repeated claim ID is its own REVIEW_BRIEF_MISMATCH. New tests: an injected adversarial controller note, an invented coverage claim, a duplicate adversarial claim-alpha. Failing first (validator at HEAD): all three published. Neutralizing the generalization (verify only): those 3 tests fail. M1: lib/review-omissions.mjs holds one rule for reconciliation and publication. reconcileLedger (and the reconcile CLI's stdout) now returns gaps: one material REVIEW_DISPOSITION_OMITTED gap per omitting required review (semantic, adversarial, coverage), naming the claim and the review's exact waveId/laneId. Contested and unsupported claims are exempt (already characterized). validate-packet rejects a packet missing such a gap (MISSING_REVIEW_OMISSION_GAP); the gap forces partial. Review Downgrades lists the claim as "not reviewed (no disposition)". Reviewers may still omit; no accepted review is rerun. Fixture changes: the two-source fixture records the helper's gaps and derives status from material gaps. createPacketFixture and the fake run now have every required review dispose of claim-2 (adversarial challenged, so it stays contested); before, the production reconciler silently left claim-2 unresolved and unreviewed in "complete" packets. Three integrity tests swap a pushed claim-2 disposition for an edit, and two omission tests use the helper's gaps instead of hand-written ones. Failing first (validator at HEAD): omitted-disposition and brief-excluded complete packets published. With the renderer at HEAD, the named-partial test fails. Neutralized: validator gap check (2 fail), renderer branch (1 fail), reconciler gaps (3 fail). M2: new test, a second claim-alpha with forged evidence on an injected source-3. Neutralizing only the duplicate-ID clause fails it and the duplicate adversarial test (2 fail). Docs: packet-contract.md, SKILL.md step 6, recon docs page. Restored suite 356/356. No version bumps. * chore(oat): record p03-t07 before re-review * chore(oat): record p03 round-3 review and stop at the review cap * chore(oat): record the operator extension and add p03-t08 * fix(p03-t08): close p03 round-3 review findings H1: lib/review-binding.mjs owns the request projection (projectReviewScope, projectReviewQuestions) and each mode's fixed excludedInputs. The generator uses them, and validateReviewBindings requires reviewBriefRequestBinds for every brief type (REVIEW_BRIEF_MISMATCH). The schema now fixes brief id (a lowercase slug, at most 64 chars) and createdAt (UTC ISO-8601). Brief-field audit (every field a reviewer sees): - kind, schemaVersion, closed top-level field set: fixed by the schema - id: lowercase slug <= 64 (new); createdAt: UTC ISO-8601 (new) - runId: bound to the ledger run; mode: bound to the review kind - excludedInputs: bound to the mode's fixed declaration (new) - scope.included/excluded, questions (adversary, coverage): bound to the manifest request projection (new); absent from verify (closed schema) - claims / provisionalStatements: each a distinct, exactly projected ledger claim (p03-t06/t07); entry keys closed - verify evidence (id, sourceId, displayExcerpt, locator): bound to ledger evidence; sources: bound to the manifest projection and claim union Failing first (validator at HEAD): 7 injection tests published (adversarial and coverage questions, an object question, scope.included, scope.excluded, verify and adversarial excludedInputs). With contracts at HEAD the id and createdAt notes published. Neutralizing reviewBriefRequestBinds: those 7 fail. The packet-contract.md sentence now lists each field it covers. M1: the omission exemption is read from the reviews (any rejected, or an adversarial challenged disposition), never from the published status. Failing first (omissions module at HEAD): an omitted claim relabeled contested or unsupported published with no gap. A challenged claim still needs no gap (control). Neutralizing the exemption: 2 tests fail. M2: one table-driven test per gap-match clause (code, material, wave, lane, claim IDs), plus a superset-claimIds control. Neutralizing each clause alone fails exactly its own case (1 each, 5 runs). Fixture: createPacketFixture's brief request now mirrors its manifest request (excludedScope); two-source fixture gains dispositionOverrides, editReconciled, and editOmissionGaps probe hooks. Suite 375/375. No version bumps; routing predicates unchanged; briefs stay blind. * chore(oat): record p03-t08 * chore(oat): record the p03 complexity review and add p03-t09 * refactor(p03-t09): rebuild recon briefs to check integrity and drop omission gaps Brief integrity is one check. validate-packet rebuilds each brief with the production createReviewBrief (its id, mode, createdAt, the manifest, the prior ledger, and the claim IDs the brief lists) and compares canonical JSON; a throw or a difference is REVIEW_BRIEF_MISMATCH naming the first differing top-level key. Dispositions must then be unique members of the brief's claims. Confirmed first: every brief type, contradiction-resolution included, is created from the prior ledger (fixtures, fake run, SKILL), so no prior-over-current overlay is needed; priorLedger ?? ledger is used. Deleted: the duplicate-ID clause, the entry loop, reviewBriefBindsClaim, briefSourceIds, briefSourcesBind, safeProjectReviewSources, reviewBriefProjection, reviewBriefRequestBinds. Kept: the projection allowlists and generator projections, and the id/createdAt schema limits. Omission-gap protocol deleted (operator decision): review-omissions.mjs, reconcileLedger's gaps return and manifest parameter, the MISSING_REVIEW_OMISSION_GAP check, the SKILL.md Step 6 duty, the contract paragraph, the docs sentence, fixture gap plumbing, and the t07/t08 omission tests. An omitted disposition behaves like uncertain: the claim stays unresolved, the packet is not forced partial, and Review Downgrades lists it as not reviewed (small renderer helper). Claim-2 fixture dispositions stay. classifyUnresolvedIssue now returns {scope, claimIds} or {error} (unresolvedIssueDiagnostic folded in; same diagnostics). The reconciler, validator, and renderer read one affirmingDispositionByReviewKind table. Tests: one table-driven tamper test (17 rows: statement, excerpt, locator, descriptor, unprojected descriptor field, injected claim on an existing and on an injected source, duplicate ID with a forged source, adversarial note and duplicate, invented coverage claim, questions x2, scope, excludedInputs, id, createdAt), each REVIEW_BRIEF_MISMATCH. Neutralizing the rebuild-and-compare check fails all 17 rows. coverage-publication test deleted (its precondition moved into end-to-end test 1); the two global-issue tests merged. KD6/KD7 controls and the end-to-end test pass. Net delta: 15 files, +389 / -1061 (-672); scripts alone +172 / -399. Suite 357/357. No version bumps; routing predicates unchanged. * chore(oat): record p03-t09 * chore(oat): record p03 review artifact * chore(oat): record gate review in project log * chore(oat): receive p03 gate finding and add p03-t10 * perf(p03-t10): rebuild each recon brief once per validation pass validateAssurance called checkReviewBrief for every verified claim, so each brief was regenerated, validated, and hashed N times (gate M1). A per-pass createBriefChecker now memoizes the rebuild-and-compare result by the exact brief reference (path and digest) plus review kind; validateReviewBindings and validateAssurance share it, and it lives only as long as one compileValidatedRun call (no module-level cache). Every guard is kept: exact brief reference, mode/run binding, rebuild-and-compare, disposition uniqueness and membership, and the 17-row tamper table. compileValidatedRun gains an optional rebuildReviewBrief seam (default: createReviewBrief) so a test can count rebuilds. Failing first: the new test counted 12 rebuilds on the two-source packet (3 bindings + 3 briefs x 3 verified claims) against the expected 3. Neutralizing the memo (always recompute) fails it again with 12. Gate scaling probe, validate time at 100/400/800 claims (local): - before this commit: 102 / 1111 / 4402 ms - after: 15 / 42 / 97 ms (pre-p03-t09 validator: 29 / 89 / 234 ms) All packets valid. Suite 358/358. No version bumps. * chore(oat): bookkeeping after p03 gate pass * fix(p04-t01): recompute next's exit-gate fingerprint with v2 exclusions oat-project-next 5.0 recomputed every stored fingerprint with only the state-carrier exclusion, so a fresh effective-delta-v2 generation read as stale after any .oat/projects or .oat/repo commit. It now recomputes v2 with the implement skill's three literal exclusions and v1 with v1 rules, and points at completion-and-closeout.md Step 14 where the algorithm lives. A contract test extracts the v2 pathspec sentence from both skills and asserts the sets are identical, and pins next's v1 rule to the carrier only. Bump oat-project-next 1.1.3 -> 1.1.4 and its pins. Failing-first: new test failed before the edit ("Expected one effective-delta-v2 exclusion pathspec sentence, found 0"). Neutralize-and-restore: dropping .oat/repo from next's v2 sentence fails the test (2 vs 3 pathspecs); swapping next's v1 pathspec for .oat/projects fails it too; both restored, test passes. Verification: vitest src/commands/init/tools/shared src/validation (isolated HOME) 902/902 pass; check:skill-bumps exit 0. * feat(p04-t02): add oat project closeout-check and guard complete-state New read-only `oat project closeout-check <path> [--autonomous] [--json]` evaluates the closeout invariant from state.md on disk. A snapshot is required when the effective layered workflow.postImplementSequence is set, when the run is autonomous (--autonomous or OAT_AUTONOMOUS=1 via an injectable env), or when oat_workflow_mode is lite. Once a snapshot exists its record is authoritative and config is never consulted. It reports status, the missing invariant, route oat-project-implement, and the next owner (stored step and skill, or approval); incomplete exits 1. complete-state takes --autonomous, runs the same check before any write, and refuses with the same message. Legacy projects recover through oat-project-implement; no override flag. Failing-first: complete-state refusal tests (cases 1 configured/flag/lite, 2, 3, 6, 7 and --autonomous) failed 8/16 before the guard existed. Neutralize-and-restore, one clause each (all restored, 49/49 green): - env read -> false: case 7 fails in closeout-check and complete-state - complete-state guard skipped: cases 1, 2, 3, 6, 7 and the trace fail - configured input ignored: configured cases and the trace fail - lite input ignored: lite cases fail - stored order -> canonical order: disk transition trace fails - approval transition removed: case 3 and the trace fail - terminal status rule removed: case 4c and the trace fail - malformed snapshot -> complete: all 8 malformed variants fail - config consulted with a snapshot: snapshot-authority tests fail Transition trace: configured and snapshot-absent state on disk, stored order document, summary, pr; a fresh invocation after every Step 15 write; interruption (summary.md in_progress) and a failed boundary both resume at summary; complete-state refuses until status complete, then succeeds. Probe (branch build): recon-rework complete; claude-effort-levels malformed (hand-written status: pending, approval: null); this project snapshot_missing. Verification: cli build exit 0; vitest src/commands/project and help-snapshots (isolated HOME) 1313/1313; oat-docs check exit 0; cli check and type-check exit 0. * feat(p04-t03): route terminal closeout through the closeout check Every terminal consumer now runs `oat project closeout-check` (adding --autonomous under OAT_AUTONOMOUS=1) and routes an incomplete closeout to oat-project-implement with the reported invariant and next owner: - implement Step 15 commits the snapshot and checks it before the first sequence child; a missing or malformed snapshot dispatches nothing. - implement Step 16 runs the check before any completion write and continues only on complete or not_required. - next 5.1 runs the check before every later route, so a configured, autonomous, or lite closeout with no snapshot now routes to implement. - complete adds Step 1.5 before its upfront questions and any write, and re-checks before complete-state, which now gets --autonomous too. oat-project-autonomous and its gate inventory are unchanged: it completes only through implement and complete. The control-plane recommender and the state dashboard stay out of scope and rely on complete-state's refusal. Bump oat-project-complete 1.7.13 -> 1.7.14 and its pins; implement and next are already bumped on this branch. Named-skill matrix rows reclassified. Failing-first: the four new contract tests failed before the edits (missing markers in implement Steps 15/16, next 5.1, complete Step 1.5). Neutralize-and-restore, one consumer each (all restored, green): - implement Step 15 pre-dispatch check removed: its test fails - implement Step 16 check removed: its test fails - next 5.1 check removed: its test fails - complete Step 1.5 gate removed: its test fails - complete Step 5 re-check removed: its test fails Verification: vitest src/commands/init/tools/shared src/validation (isolated HOME) 906/906; node --test oat-project-implement 37/37; oat:validate-skills exit 0; check:skill-bumps exit 0. * feat(p04-t04): add operator-only exit-gate waivers DR-260927-operator-waiver-for-test-only: no automatic test-only exception. oat_implement_exit_gate gains an append-only `waivers` list; each entry records waived_by, reason, from_commit..to_commit (Git range), the covered_fingerprint at to_commit in the generation's own version, and a UTC waived_at. A waiver never rewrites reviewed_head, the implementation or freshness fingerprints, freshness_head, or an earlier waiver, and never revives a generation already persisted stale. It is written only on an explicit operator instruction, never inferred or self-issued; under OAT_AUTONOMOUS=1 every waiver write is refused and the run stops. A waived generation reads allowed only while nothing substantive lands after the covered range; v1 and v2 share one rule and differ only in which paths are substantive. Malformed or unverifiable waivers fail closed as stale. next 5.0 validates waivers read-only; summary and pr-final list every waiver; the state template and three docs pages describe the field. Bump oat-project-pr-final 1.6.6 -> 1.6.7 and pins; summary, implement, and next are already bumped on this branch. Failing-first: the six new waiver tests failed before the edit (missing Operator waivers section). Neutralize-and-restore, one clause each (all restored, green): - operator-instruction clause removed: issuance test fails - OAT_AUTONOMOUS refusal removed: issuance test fails - append-only clause removed: record test fails - malformed fail-closed weakened to "ignored": v1/v2 rule test fails - later-substantive-stale clause removed: v1/v2 rule test fails - per-version covered fingerprint removed: v1/v2 rule test fails - pr-final listing made optional: summary/PR test fails - reason dropped from the skill record: record test and the v1 and v2 executable checks fail (the check reads its field set from the skill) - executable check without waiver validation: v1 and v2 checks fail - executable check covering past to_commit: v1 and v2 checks fail Verification: vitest src/commands/init/tools/shared, src/commands/project, src/validation (isolated HOME) 2162/2162; oat:validate-skills exit 0; check:skill-bumps exit 0; oat-docs check exit 0. * chore(oat): record p04 task ledger before review * chore(pjm): run a complexity review when review or gate budgets run out Adds BL-261001-run-a-complexity-review-when as the first slice and gives BL-260818 the complexity-review input and a fourth disposition, simplify. * chore(oat): receive p04 review and add p04-t05 * fix(p04-t05): close p04 review findings Review: reviews/archived/p04-review-2026-10-01T174412Z.md (M1, L1-L3). M1: the waiver is now reachable. Operator waivers gain a stale-boundary offer: before an allowed qualified generation that is stale only through uncovered descendants is persisted stale, an interactive run lists the commits and asks the operator to waive that range or start a new gate run; stale is persisted only when the operator declines. Under OAT_AUTONOMOUS=1 no waiver is offered or issued: persist stale and start a new gate run. Other stale causes get no offer. An operator may also record a waiver before resuming implement. Step 14 item 1 and the Step 15 freshness re-check route through the offer; next 5.0 mirrors it in its stale note. Docs (implementation-execution, workflow-gates) describe both paths. L1: transition-trace step 3 wrote summary.md, which the check never reads; removed. The failed boundary is now the interruption step. L2: the BL-260829 comment moved back above its own describe. L3: an approval_pending nextOwner now carries both writes, `approval: approved` and `approval: not_required`; the message names both; Step 15 says nextOwner is a step to dispatch or, with an empty pre_approval, the approval boundary to record. Autonomy coverage maps the three new L3 lines to NG; named-skill matrix rows added for the two new sentences. Failing-first: 3 new tests failed before the edits (offer clauses, Step 15 nextOwner wording, case 3c empty pre_approval). Neutralize-and-restore (each restored, green): - interactive offer removed: offer test fails - autonomous never-offer clause removed: offer test fails - Step 14 item 1 bypasses the offer: offer test fails - next mirror removed: offer test fails - approval writes dropped from nextOwner: case 3c fails Verification: cli build exit 0; vitest src/commands/project, src/commands/init/tools/shared, src/validation (isolated HOME) 2165/2165; check:skill-bumps exit 0; cli check, type-check, validate-skills exit 0. No version bumps. * chore(oat): bookkeeping after p04 pass * chore(oat): record p04 review artifact * chore(oat): record gate review in project log * chore(oat): receive p04 gate findings and add p04-t06 * fix(p04-t06): close p04 gate findings Gate: reviews/archived/p04-review-2026-10-01T180727Z.md (M1, L1). M1: a failed snapshot chose `prePending ?? postPending` before checking approval, so with pre-approval work done, approval pending, and retro stored post-approval it named retro. The failed branch now picks its owner in approval-aware order: pending pre-approval work, the pending approval boundary (both writes named), post-approval work, then status repair. A post-approval step is never named while approval is pending. The approval owner is shared with the approval_pending branch. L1: implementation-execution and lifecycle docs now say an autonomous run that finds a stale generation persists stale and starts a new gate run; autonomy refuses only an attempted waiver write. Failing-first: the gate's exact fixture, added as a closeout-check command-boundary regression, failed before the fix (nextOwner was the post_approval retro step). Two ordering companions (pre-approval work first; post-approval work once approval is recorded) pin the other legs. Neutralize-and-restore: putting post-approval work ahead of the approval boundary again fails the fixture test; restored, 32/32 green. Branch-CLI probe with the gate's fixture and command: sequence_failed, nextOwner kind approval with both writes, exit 1. Verification: cli build exit 0; vitest src/commands/project (isolated HOME) 1254/1254; oat-docs check exit 0; cli check and type-check exit 0. No version bumps. * chore(oat): bookkeeping after p04 gate pass * fix(p05-t01): never overwrite a CLAUDE.md that AGENTS.md links to Under a shim strategy with --force, instructions sync overwrote a CLAUDE.md that an AGENTS.md resolves to, destroying the only copy of the instructions. Planning now keeps such a file and reports it as a skip. The resolves-to check (hard link, or a symlink chain ending at CLAUDE.md) is extracted from inspectManagedShim and shared with the new findAgentsResolvingTo helper; a CLAUDE.md that is itself a symlink counts only when an AGENTS.md chain passes through it, so the ordinary CLAUDE.md -> AGENTS.md shim is still replaced. Single planning-time guard; no apply-time re-check. Failing-first: 5 new integration tests failed before the guard (pointer and symlink --force on AGENTS.md -> CLAUDE.md, a symlink chain, a hard link, and a cross-directory link); the non-linked control passed before and after. Neutralize-and-restore: making keepClaudeFilesAgentsResolveTo return its input unchanged failed the same 5 tests (143 passed); restored, 148/148 pass. Verify: HOME=$(mktemp -d) vitest run src/commands/instructions -> exit 0. * refactor(p05-t02): make dispatch record validate-only Per DR-260927-dispatch-record-validates, remove dispatch-record persistence and keep the validate-only command plus its schema modules for the managed Claude validation path. - record.ts: delete the <project>/dispatch/ journal writer, writer lock, revision naming, fallback-claim publication, trigger and related-record reads; recordProjectDispatch now takes { input } and returns { status: 'validated-only', record, runtimeIdentity }. - index.ts: drop --project, the persisted status, and the <project> redaction label. - record.test.ts: prune journal tests (61 -> 49 cases); move the journal-byte redaction assertions (absolute, assignment-form, nested evidence, rollout, Claude answer) onto the command's validate-only JSON output via validatedOutput(); add a fallback-link refusal case. - Skills and docs: remove persistence wording from oat-dispatch-subagents (SKILL.md, record-schema.md, managed example), oat-project-dispatch-subagents, review-provide, review-provide-remote, plan-writing, implement dispatch-and-dry-run.md, CLI reference, evidence-layers, orchestration-model (Journal participant -> run record), implementation-execution, scope-and-surface. - Bumps: oat-project-dispatch-subagents 1.1.7, oat-project-review-provide 1.5.12, oat-project-review-provide-remote 1.1.9, oat-project-plan-writing 1.2.35; pins updated. Failing-first: the rewritten skills.test.ts case (persistence wording absent) failed on oat-dispatch-subagents/SKILL.md before the doc edits. Neutralize-and-restore: appending the old opt-in sentence to plan-writing SKILL.md failed that case; re-adding a --project option failed the help-snapshot case; both restored. Probe: dist project dispatch record --project x --event-file - -> unknown option, exit 1. Search from the plan: no output, rg exit 1. Verify: build 0; vitest dispatch+help+validation 710/710; test:smoke 163/163; test:skills 689/689; oat:validate-skills 0. * feat(p05-t03): let test-only package changes skip the lockstep bump Per DR-260927-test-only-paths-skip, each public package's versionPolicyIgnorePatterns now lists exactly the test paths its tsconfig.json excludes from dist: cli src/**/*.test.ts and src/**/__tests__/** (beside assets/**), control-plane **/*.test.ts and **/*.spec.ts, docs-config and docs-transforms src/**/*.test.ts, docs-theme none. AGENTS.md Package Management states the rule. Tests: a contract test reads each package's tsconfig exclude (minus node_modules, dist) and requires it to equal the ignore patterns; path cases for every package; release-utils and runVersionBumpCheck cases for a cli test-only change (no bump) and a non-test src change (bump required, the negative control). Failing-first: 4 new cases failed before the pattern change (contract equality, path cases, test-only diff, test-only check). Neutralize-and-restore: widening the cli pattern to src/** failed 5 cases including the negative control; restored, 77/77 pass. Real probe in a scratch worktree off origin/main (contract change applied uncommitted): test-only commit to packages/cli/src/release/release-utils.test.ts -> release:check-versions exit 0 ("no public package changes"); then a commit to packages/cli/src/index.ts -> exit 1 ("Changed packages: @open-agent-toolkit/cli"). Worktree removed with git worktree remove --force. Deviation: the probe ran pnpm install --frozen-lockfile instead of worktree:init; see the phase report. * feat(p05-t04): report YAML errors with location and check key types parseSkillFrontmatter now uses a LineCounter and returns the first parser error's position (or the non-mapping metadata node's position) as a problem. skill-frontmatter-unreadable appends it as a SKILL.md line and column, e.g. "...: line 3, column 14: Nested mappings are not allowed in compact mappings". New findSkillFrontmatterTypeProblems checks parsed types of name, description, allowed-tools (string) and disable-model-invocation, user-invocable (boolean); the oat-* loop reports skill-frontmatter-type findings beside the existing semantic checks, which are unchanged. Failing-first: 4 new cases failed before the change (bare-colon description, name: 123, user-invocable: "true", metadata: [1.0.0]); the valid-fixture control passed. Five existing malformed-frontmatter expectations gained the location suffix. Neutralize-and-restore: dropping the problem from the finding and skipping type problems failed the same 4 cases; restored, 622/622 pass. Verify: HOME=$(mktemp -d) vitest src/validation src/commands/shared -> 622 pass; pnpm oat:validate-skills -> 65 skills OK. * fix(p05-t05): route quick-mode discovery to quick-start The status/list recommender and the state dashboard routed quick-mode discovery through oat-project-plan, a two-hop route that disagreed with the oat-project-next and oat-project-progress skill tables. Changed exactly: router.ts discovery:in_progress:2 and discovery:complete:1, and generate.ts quick:discovery:complete, to oat-project-quick-start. Router discovery:in_progress:3 and the dashboard's quick:discovery:in_progress stay at oat-project-discover, now pinned by tests. Quick plan:in_progress routing is untouched (BL-261001-route-quick-mode-plan). Mechanical propagation (outside the declared files): split/__tests__/run.test.ts pinned the old route for a seeded quick child's project status; it now expects quick-start. Failing-first: 3 router cases and 1 dashboard case failed after the pins were flipped and before the route change. Neutralize-and-restore: pointing router discovery:in_progress:3 at quick-start failed "keeps quick discovery tier 3 in discovery"; pointing the dashboard's quick:discovery:in_progress at quick-start failed the dashboard case; both restored. Verify: control-plane vitest src/recommender 49/49; HOME=$(mktemp -d) cli vitest src/commands/state 31/31; src/commands/project/split 39/39. * fix(p05-t02): prune dispatch-record persistence tests missed in p05-t02 Recovery event bw3-p05-recovery-1 (attempt 1/10, phase-standing) for original commit a92f397fd. p05-t02 removed --project and the persisted status, but three test files outside its verification set still exercised persistence and failed: commands.integration.test.ts (2 cases), e2e/workflow.test.ts (2 cases), and init/tools/shared/review-skill-contracts.test.ts (pinned the removed opt-in wording). Bounded, test-only correction: run the command without --project, expect status validated-only, assert on stdout instead of journal bytes, and assert the persistence wording is absent. No production code changes. Verify (pre-commit): focused vitest src/commands/project src/e2e src/commands/commands.integration.test.ts 1362/1362; phase vitest src/commands src/validation src/release 5979/5979; whole CLI suite 8012/8012. Ledger: state.md p05 pending_attempt marked completed. * docs(p05-t06): narrow the packs inventory redaction claim troubleshooting.md claimed "Reported project and home paths remain redacted" for the packs:inventory diagnostic. The code redacts less, so the entry now says what it does: the project root (to a relative path or .) and the home root (to ~) are replaced only when that scope is part of the run, by literal replacement of the exact root path, and a path outside those roots, such as a global bundle path in an assets error, stays absolute. Code paths (packages/cli/src/commands): - status/index.ts unavailablePackReport (~204-240): replaceAll over projectRoot/userRoot pairs built only from the run's scopeRoots; collectPackReport (~708-713) builds roots. - tools/shared/format-pack-inventory.ts redactPackText (~23-62): same literal replacement for unavailablePackEvidence reasons. - doctor/index.ts (~1098-1103 roots, ~1186-1213 packs:inventory check): message is the evidence diagnostic detail from redactPackText. BL-260903-verify-the-packs-inventory: placeholder acceptance criteria filled with this outcome. No new test (docs-only, per plan). Verify: pnpm --filter oat-docs check -> 0. * chore(oat): record p05 task ledger before review * chore(oat): receive p05 review and add p05-t07 * chore(pjm): file wave 3 follow-ups for bundle-assets and oat-wrap-up * chore(oat): fix the bundle-assets backlog reference * fix(p05-t07): close p05 review findings M1 (tools/release/release-utils.ts): findChangedWorkspaceDirsFromPaths judged a path under a dependency root by the dependent's ignore patterns, so a control-plane or docs-transforms test-only change still marked packages/cli or packages/docs-config. Each path inside a public package is now judged by the patterns of the package that owns it; paths outside every package (additional roots) keep the contract's own patterns. - Failing-first: with the real dependency map, the control-plane and docs-transforms test-only cases failed; the non-test negative control (router.ts -> cli + control-plane; docs-transforms index.ts -> docs-config + docs-transforms) passed before and after. - Neutralize-and-restore: dropping the owner lookup failed the 2 test-only cases; restored, src/release 80/80. M2: delete journal-only machinery with no production caller. - fs/io.ts: withContainedWriterLock, journalErrorCode, redactedFsError, publishContainedJsonRevision and their private helpers; io.test.ts lock/publication cases and the journal race mocks. - absolute-paths.ts: assertJournalIdentityHasNoAbsolutePath; journal docstring reworded (also two journal comments in generic-dispatch-record.ts). - oat-dispatch-record.ts: triggerRecord/relatedRecords inputs, the fallback-claim and fallback-link event kinds and branches, and the fallback control-field lists and digest only they used, with their tests. The record keeps fallbackClaim: null and a not-applicable fallback, so the validated output shape is unchanged. - Failing-first: record.test.ts now expects fallback-link and fallback-claim to be refused as unknown kinds (Invalid discriminator value); it failed before (got "Fallback requires the rejected trigger record"). The review's rg for the deleted names went from 60 matches to none. - Unchanged validate-only/managed path: branch CLI on managed-claude-example.json implementer and reviewer inputs -> exit 0, status validated-only, keys record/runtimeIdentity/status, payload.variant intact; managed and dispatch tests pass. Verify: cli build 0; vitest src/release src/fs src/providers src/commands/project/dispatch 1102/1102; test:smoke 163/163; cli type-check 0; pnpm lint 0 (0 cached); whole CLI suite 7954/7954; pnpm check 0. No skill text changed, so no version bumps. * chore(oat): bookkeeping after p05 pass * chore(oat): record p05 review artifact * chore(oat): record gate review in project log * chore(oat): receive p05 gate finding and add p05-t08 * fix(p05-t08): compare inodes through symlinked AGENTS.md Gate M1: the --force preservation check compared realpaths when AGENTS.md was a symlink, so pkg/AGENTS.md -> ../alias.md, with alias.md hard-linked to CLAUDE.md, was missed and --force replaced the protected CLAUDE.md. agentsResolvesToClaude now resolves a symlinked AGENTS.md to its endpoint and compares that endpoint's device and inode with the regular CLAUDE.md. Read failures still fail closed (unable to resolve AGENTS.md; kept). Callers pass only a regular CLAUDE.md, so the ordinary CLAUDE.md -> AGENTS.md shim is still overwritten. The unused claudePath parameter is dropped. Failing-first: the new command-boundary regression (root AGENTS.md independent, alias.md a hard link of CLAUDE.md, pkg/AGENTS.md -> ../alias.md, sync --force) failed for pointer, symlink and copy; each now skips the CLAUDE.md action, and its content and inode are unchanged. Direct-symlink, chain, hard-link and ordinary-overwrite controls still pass. Neutralize-and-restore: making the endpoint comparison always false failed the 3 new cases and the 5 symlink-based keep cases (8 failed); restored, 151/151 pass. Verify: HOME=$(mktemp -d) vitest run src/commands/instructions -> exit 0, 151/151; cli check and type-check 0. * chore(oat): bookkeeping after p05 gate pass * chore(p06-t01): bump lockstep public packages to 0.3.11 Follows the version files #334 (98d1d5246) changed: the five public package.json files and packages/cli/assets/public-package-versions.json. The sync manifest's oatVersion is a sync stamp, not a lockstep version, and is left unchanged. Verified: release:check-versions exit 0 against origin/main 98d1d5246; release:validate exit 0 (5 packages at 0.3.11); scoped release and bundle tests pass. * chore(p06-t02): archive the backlog items shipped in wave 3 Archived 13 items with the branch CLI (oat backlog archive, all exit 0). BL-260909 cites DR-260927-dispatch-record-validates; BL-261001-recover-recon-lanes-after cites discovery Key Decision 8; BL-260829 cites the p01/p02 Step 7a reviews (de9c98848, c129e82ab) under oat-project-implement 2.3.14 (origin/main 8f6d5b1d2). BL-260806 stays open until closeout. The BL-260829 archive rewrote two inbound references. Adds the Wave 3 curated overview note; regenerate-index is idempotent; pjm doctor reports only its pre-existing warning. * chore(oat): record p06 task ledger before review * chore(oat): close p06 review findings and write the PR hand-off * chore(oat): record p06 review artifact * chore(oat): record gate review in project log * chore(oat): prepare final implementation closeout * chore(oat): record the final review and close its findings * chore(oat): persist exit gate generation 1 intent * chore(oat): record exit gate generation 1 acceptance * chore(oat): record final review artifact * chore(oat): record gate review in project log * chore(oat): persist exit gate generation 1 result and receive intent * chore(oat): receive exit gate generation 1 review * chore(oat): exit gate generation 1 allowed * chore(oat): persist the post-implement sequence snapshot * docs: generate summary for backlog-wave-3 Promote five key decisions to reference/decisions (DR-261001-*); roll up 20 project-log entries. * chore(oat): record the summary closeout step * chore(pjm): record backlog wave 3 in current state and roadmap Add the CLI 0.3.11 entry to current-state.md and drop the closed BL-260829-order-phase-bookkeeping-before bullet from the roadmap's Next list. * docs(backlog-wave-3): update documentation from project artifacts - contributing/code.md: test-only changes skip the lockstep bump (replaces the stale claim) - troubleshooting: older oat rejects template resolve and closeout-check; need 0.3.11+ - lifecycle: oat-project-complete runs closeout-check before any completion write - instruction-sync: --force keeps a CLAUDE.md that an AGENTS.md resolves to - recon: brief rebuild-and-compare, scoped unresolvedIssues, coverage gaps downgrade - add-docs-to-a-repo: the Fumadocs scaffold prebuild does not run nav sync --check * chore(backlog-wave-3): mark docs updated * chore(oat): record the document closeout step * chore(oat): archive residual plan reviews and mark the final PR ready Archive the two fixes_completed plan artifact-review files under reviews/archived/ and rewrite their plan.md ledger and implementation.md references (pr-final Step 0.5). Set oat_pr_status to ready now that the final PR artifact exists. * chore(oat): record final PR metadata Final PR opened as #336 (https://github.com/voxmedia/open-agent-toolkit/pull/336). Set oat_pr_status to open, oat_pr_url, and oat_phase_status to pr_open, and route the next milestone to oat-project-revise or either oat-project-complete ordering. * chore(oat): record the pr closeout step * docs(oat): build backlog-wave-3 project recap Unattended project-recap run e33c16db-fd14-47c7-9c4e-2cdccfef6d9b, outcome built. Host verify rung: headless Chrome at 320, 768, and 1440 pixels; all browser-free checks pass. * docs(oat): record the recap outcome in the backlog-wave-3 summary * docs(oat): rebuild backlog-wave-3 project recap Rebuild after d4fe765f8 added the Explainer Outcome section to summary.md, an approved input. Unattended project-recap run 810e34a7-ca5d-4437-a88a-e1cd02ec37a7 replaces e33c16db in place. Outcome built; host verify rung at 320, 768, and 1440 pixels; all browser-free checks pass. * chore(oat): closeout awaiting final approval * chore(oat): record autonomous final approval * chore(oat): closeout sequence complete * chore(oat): mark implementation complete * chore(oat): record the BL-260806 live closeout trace * chore(pjm): archive BL-260806 with the wave 3 closeout trace * docs(oat): rebuild backlog-wave-3 project recap for completion Rebuild after implementation.md gained the final-approval record and the BL-260806 closeout trace. Unattended project-recap run b2bc7235-9374-4a84-84a4-3abbc6b2ebe2 replaces 810e34a7 in place. Outcome built; host verify rung at 320, 768, and 1440 pixels; all browser-free checks pass. * chore(oat): complete project lifecycle for backlog-wave-3
Adds selectable GPT-6.1 Sol pins for Codex and Sonnet 5.5 pins for Claude and Cursor, with low, medium, high, xhigh, and max efforts. The bundled recommendation now uses Sol 6.1 in Codex High/Frontier and Sonnet 5.5 medium in Claude Economy. Earlier Sol pins remain supported.
Cursor 3.22.12 silently fell back for every Sonnet 5.5 bracket selector, but resolved all five exact model IDs correctly. OAT now permits explicitly registered exact-ID pins with matching native probe evidence. Shared validation rejects missing or mismatched records before sync or rendering. The PR retains redacted native subjects and controls, updates the pin-verification runbook and model guidance, and regenerates provider roles. Public packages move together to 0.3.10.
Validation:
Protects against: silently approving a Cursor selector that falls back to another model; structural validation alone cannot detect provider resolution; covered by captured native lifecycle fixtures and mapping consistency tests.