Skip to content

feat(dispatch): add Sonnet 5.5 and GPT-6.1 Sol pins - #334

Merged
tkstang merged 1 commit into
mainfrom
feat/sonnet55-sol61-pins
Oct 1, 2026
Merged

tkstang merged 1 commit into
mainfrom
feat/sonnet55-sol61-pins

Conversation

@tkstang

@tkstang tkstang commented Oct 1, 2026

Copy link
Copy Markdown

Adds selectable GPT-6.1 Sol pins for Codex and Sonnet 5.5 pins for Claude and Cursor, with low, medium, high, xhigh, and max efforts. The bundled recommendation now uses Sol 6.1 in Codex High/Frontier and Sonnet 5.5 medium in Claude Economy. Earlier Sol pins remain supported.

Cursor 3.22.12 silently fell back for every Sonnet 5.5 bracket selector, but resolved all five exact model IDs correctly. OAT now permits explicitly registered exact-ID pins with matching native probe evidence. Shared validation rejects missing or mismatched records before sync or rendering. The PR retains redacted native subjects and controls, updates the pin-verification runbook and model guidance, and regenerates provider roles. Public packages move together to 0.3.10.

Validation:

  • All required checks passed: check, type checking, forced workspace tests, smoke/skill/script tests, build, skill/version gates, release validation, and docs build; lint and format also passed.
  • Cursor native verification: 16 launches and 64 correlated events across bracket and exact-ID rounds, with positive and negative controls and no Task model overrides.
  • Removing the exact-ID evidence guard made four rejection tests fail; restoring it returned the focused suite to green.
  • Branch CLI sync is idempotent; a scratch repository generated all ten Claude Sonnet 5.5 role variants across five efforts.
  • Independent review found no blocking issues.

Protects against: silently approving a Cursor selector that falls back to another model; structural validation alone cannot detect provider resolution; covered by captured native lifecycle fixtures and mapping consistency tests.

@cursor

cursor Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Changes default subagent models and Cursor pin rules for lifecycle dispatch; misconfigured selectors could silently fall back without the new probe-backed exact-ID gates.

Overview
This PR refreshes OAT’s model dispatch catalog for an October 2026 availability pass: GPT-6.1 Sol becomes the default Codex Sol selector (efforts unchanged; older Sol pins stay supported), and Claude Sonnet 5.5 is added across Claude and Cursor with full effort rungs. The bundled dispatch matrix shifts Codex High/Frontier to gpt-6.1-sol and Claude Economy to Sonnet 5.5 medium.

For Cursor, native probes showed Sonnet 5.5 bracket selectors fall back while exact model IDs resolve correctly. The pin-mapping decision and dispatch docs now allow explicit exact-ID registry entries with per-mapping probe evidence—catalog visibility alone cannot pass through arbitrary IDs. Probe fixtures, runbook references, and skill/provider guidance are updated accordingly.

Codex gains new oat-phase-implementer and oat-reviewer agent definitions for gpt-6.1-sol at low through max, registered in .codex/config.toml. Related workflow docs, dispatch log examples, and package/sync metadata bump to 0.3.10.

Reviewed by Cursor Bugbot for commit 6c18f4f. Bugbot is set up for automated code reviews on this repo. Configure here.

@tkstang
tkstang merged commit 98d1d52 into main Oct 1, 2026
3 checks passed
tkstang added a commit that referenced this pull request Oct 1, 2026
Resolve skills.test.ts pin conflicts (main's plan-writing 1.2.34 plus wave 3 bumps) and move
oat-project-implement to 2.3.16 and oat-dispatch-subagents to 1.2.12 above main's 2.3.15/1.2.11.
tkstang added a commit that referenced this pull request Oct 1, 2026
Follows the version files #334 (98d1d52) changed: the five public package.json
files and packages/cli/assets/public-package-versions.json. The sync manifest's
oatVersion is a sync stamp, not a lockstep version, and is left unchanged.

Verified: release:check-versions exit 0 against origin/main 98d1d52;
release:validate exit 0 (5 packages at 0.3.11); scoped release and bundle tests pass.
tkstang added a commit that referenced this pull request Oct 1, 2026
… validate-only dispatch record (wave 3, lockstep 0.3.11) (#336)

* chore(pjm): set wave 3 lead items

Raise BL-260927-expose-a-scoped-template to high as the wave lead and give
BL-260718-support-fumadocs-in-oat-docs acceptance criteria: nav sync writes Fumadocs
meta.json from authored index.md Contents maps.

* chore(pjm): file recon issue #333 as three backlog items

The packet-validator contract fix joins wave 3 at high priority; Codex capacity recovery and
controller setup friction are tracked at medium and low.

* chore(pjm): scope recon recovery to a note and one retry

Mixed native and CLI continuation records move to a deferred item until the friction recurs.

* chore(pjm): file two lifecycle routing items from wave 3 recon

* chore(oat): scaffold backlog-wave-3

* chore(oat): enable all-phase gates for backlog-wave-3

* chore(oat): persist autonomous explainer intent

Also records the managed/high dispatch policy chosen for this wave.

* chore(oat): update autonomous execution learnings

* chore(oat): capture quick-start discovery for backlog-wave-3

* chore(oat): complete quick-start discovery for backlog-wave-3

* chore(oat): draft backlog-wave-3 plan for review

* chore(oat): apply plan artifact review findings

* chore(oat): update autonomous execution learnings

* chore(oat): apply plan re-review findings

* chore(oat): record plan artifact review disposition

* chore(oat): record plan review artifact

* chore(oat): record gate review in project log

* chore(oat): resolve plan gate and complexity review findings

* chore(oat): record plan review artifact

* chore(oat): record gate review in project log

* chore(oat): record plan gate attempt 2 and stop at the QS-12 boundary

* chore(oat): update quick-start artifacts for backlog-wave-3

* chore(oat): initialize backlog-wave-3 implementation

* chore(oat): add the default phase recovery ledger

* feat(p01-t01): share one template resolver in repository, user, bundle order

Move the PJM template resolver to commands/shared/template-source.ts as
resolveTemplate (repository, user, bundle) and drop "PJM" from its
messages. pjm init, backlog new and decision new call it unchanged.
Project scaffold and lite promotion now call it with
templatesRoot=<repo>/.oat/templates instead of the user-first
resolveTemplateSource, which is deleted. The bundle root may be passed
lazily so scaffold still resolves assets only when both tiers miss.

Unchanged on purpose: cleanup/project/project.ts (repository or inline)
and project/log/append.ts (bundle only). Neither copies a lifecycle
template that a user-scope install lacks.

Failing first: flipped scaffold.test.ts "uses a repo template before a
differing user template" failed before the change with
"expected '# repo-first USER-TEMPLATE...' to contain 'REPO-TEMPLATE'".
The partial-tier scaffold test also encoded user-first order; it now
proves repo plan, user discovery and bundle state in one scaffold.

Verify: HOME=$(mktemp -d) vitest run src/commands/{shared,project,pjm,
backlog,decision}: 126 files, 2365 tests passed; pnpm type-check exit 0.

* feat(p01-t02): add oat template resolve

Add `oat template resolve <name> [--output <path>]` (global --json).
It resolves through the shared repository, user, bundle resolver and
reports name, found, tier and path. path is null for the bundle tier,
so no package-manager path leaks. <name> takes plan or plan.md and
rejects separators and "..". --output copies the content, replacing an
existing file, and creates no parent directories. A miss exits 1 and
names the three tiers. Outside a git repository the repository tier is
skipped. The shared resolver gains TemplateNotFoundError (same message)
so the command can tell a miss from a read failure.

Failing first: with the help snapshots and template/resolve.test.ts
written before the command, vitest failed 3 snapshot tests (root help,
template help, template resolve help) and the resolve suite failed to
import the missing module.

Verify: build exit 0; HOME=$(mktemp -d) vitest run src/commands/template
src/commands/help-snapshots.test.ts: 77 tests passed; oat-docs check
exit 0; type-check and lint exit 0.
Branch-CLI probe (issue #296): temp git repo with no .oat/templates and
HOME holding ~/.oat/templates/plan.md -> tier user, --output copied it;
discovery -> tier bundle, path none; nope -> exit 1 found:false;
../x -> exit 1; --output into a missing dir -> exit 1.

* feat(p01-t03): route lifecycle skills through oat template resolve

Lifecycle skills copy templates with
`oat template resolve <name> --output "$PROJECT_PATH/<file>"` instead of
reading .oat/templates/, so a user-scope-only install works. Unconditional
copies are straight substitutions; fill-if-missing steps (quick-start,
import-plan plan.md) keep their condition in prose. Retro now creates
references/ first because --output creates no directories.

Skills: retro, design (spec + 3 design copies), spec, discover (state,
discovery), plan (copy + overwrite), summary, quick-start (resume, scaffold
note, discovery, design, plan), import-plan (plan, implementation),
promote-spec-driven, implement plan-and-resume, new, repo-improve,
pjm-decision. cursor-cloud-projects: templates resolve repository, user,
bundle via the CLI; skills and scripts keep user-first order, and the
staleness rationale now covers only skills and scripts. Ideas skills are
unchanged. Docs: cursor-cloud, file-locations, tool-packs, troubleshooting.

Pins: skills.test.ts quick-start/import-plan rewrite markers and every
bumped version; review-skill-contracts.test.ts quick-start resume pins.
Bumps (patch): retro 1.0.8, design 2.3.7, spec 2.0.4, discover 2.2.8,
plan 1.4.16, summary 1.5.7, quick-start 2.3.17, import-plan 1.4.19,
promote-spec-driven 1.2.4, implement 2.3.15, cursor-cloud-projects 1.1.3,
new 1.4.2, repo-improve 2.1.6, pjm-decision 1.1.2.

Failing first: after the text edits, vitest src/validation failed 16
tests (stale version pins and moved rewrite markers) and
init/tools/shared failed the quick-start resume pin; both pass now.

Verify: rg "\.oat/templates/" .agents/skills (non-test) leaves only
descriptive matches; oat:validate-skills exit 0; HOME=$(mktemp -d)
vitest run src/validation src/commands/init/tools/shared: 899 passed;
node --test oat-project-implement/tests exit 0; oat-docs check exit 0;
oxlint .agents/skills exit 0; oat sync --scope project --dry-run: no
changes.

* chore(oat): record p01 task ledger before review

* chore(oat): receive p01 review and add p01-t04

* fix(p01-t04): close p01 review findings

M1: oat-project-retro and oat-project-summary curate oat access one
subcommand at a time, so both now grant Bash(oat template:*) for their
template copy; retro also grants Bash(mkdir:*) for creating references/.
skills.test.ts pins both grants next to the Bash(oat tools:*) pin. The
other p01-t03 skills need nothing: discover and implement grant
Bash(oat:*), the rest grant bare Bash, and design/spec/plan's git-only
list predates p01 (review scoped it out). Both skills were already bumped
on this branch (1.0.8, 1.5.7), so no further bump.

The summary allowed-tools edit changed that line's autonomy prompt-site
key, so .agents/docs/autonomy-contract.md maps the new key a8983fec040c
-> NG in place of 35cb2ea1d677 -> NG (same disposition).

L1: tool-packs.md moves the oat template resolve sentence above the
fallback paragraph, so "Useful options" stays attached to oat pjm init.

Failing first: the new pin failed before any grant ("expected ... to
contain 'Bash(oat template:*)'"), then failed again on 'Bash(mkdir:*)'
with only the oat template grants added; it passes with both.

Verify: HOME=$(mktemp -d) vitest run src/validation: 391 passed;
oat:validate-skills, check:skill-bumps and oat-docs check exit 0.

* chore(oat): bookkeeping after p01 pass

* chore(oat): record p01 review artifact

* chore(oat): record gate review in project log

* chore(oat): receive p01 gate findings and add p01-t05

* fix(p01-t05): close p01 gate findings

M1: oat-project-design, oat-project-spec and oat-project-plan now run
oat template resolve, so each allowed-tools list grants
Bash(oat template:*). The p01-t04 allowlist contract test in
skills.test.ts now covers all three alongside retro and summary. The
design and plan header edits changed their autonomy prompt-site key, so
.agents/docs/autonomy-contract.md maps f9e319c9b223 -> NG in place of
5eb3949f32e1 -> NG (same disposition); spec is not an autonomy root.
All three skills already carry their branch bumps (2.3.7, 2.0.4, 1.4.16).

L1: file-locations.md, tool-packs.md and troubleshooting.md now limit
the repository, user, bundle claim to project lifecycle and PJM
templates and note that idea templates follow the ideas scope. The
resolver is unchanged.

Failing first: the extended test failed before any grant on design
("expected 'name: oat-project-design...' to contain
'Bash(oat template:*)'"). Adding grants one at a time, it then failed on
spec, then on plan, and passed once all three were granted.

Verify: HOME=$(mktemp -d) vitest run src/validation: 391 passed;
oat:validate-skills, check:skill-bumps and oat-docs check exit 0.

* chore(oat): bookkeeping after p01 gate pass

* feat(p02-t01): write Fumadocs meta.json from index.md Contents maps

`oat docs nav sync` detects the framework: mkdocs.yml keeps the MkDocs path;
a Fumadocs source.config.* writes one strict meta.json per docs directory
with an index.md. Pages follow the Contents order with no "..." rest entry,
the root lists "index" first, subfolders appear by name, and cross-folder
Contents links become Fumadocs link entries. The folder title comes from
the index.md frontmatter title, else its first H1. Existing files are
compared by meaning, so a rerun writes nothing even after reformatting.
Unlisted pages and folders are reported in human output and --json.

Failing first: 8 new tests in sync.test.ts failed before the change
(framework undefined, no meta.json, no command reporting); 6 existing passed.
Neutralize and restore: forcing changed=true failed the rerun and
no-changes tests; naming cross-folder pages instead of linking failed the
nested-fixture test; both pass after restore.
Verify: vitest run src/commands/docs, 14 files, 179 tests pass.

* docs(p02-t02): describe Fumadocs nav sync

Help text and the docs pages now describe both frameworks: nav sync writes
mkdocs.yml nav for MkDocs and strict, committed meta.json files for Fumadocs,
reports unlisted pages, and writes cross-folder Contents links as link
entries. Adds a Fumadocs Nav Sync section to the docs index contract.
Also updates reference/file-locations.md and reference/index.md, which
carried the same MkDocs-only wording, and regenerates the tracked
apps/oat-docs/index.md manifest for the changed page descriptions.

Failing first: the docs nav sync --help inline snapshot failed against the
new description and option text before the snapshot update.
Verify: vitest help-snapshots + src/commands/docs (15 files, 241 tests)
and pnpm --filter oat-docs check pass.

* docs(p02-t03): generate Fumadocs navigation for apps/oat-docs

Ran the branch CLI `docs nav sync --target-dir apps/oat-docs`: it wrote 11
meta.json files and reported no unlisted pages, so all 70 pages stay in the
sidebar and no Contents map needed changes. workflows/skills links into
contributing/ and docs-tooling/ became link entries.

Failing first: before this commit apps/oat-docs had no meta.json, so the
sidebar used Fumadocs' default file-tree order instead of Contents order.
Verify: a second run after `oxfmt --write apps/oat-docs/docs` reported
written [] and unlisted []. pnpm build:docs passed uncached (0 of 6
cached); the built sidebar order for the root, workflows/projects and
contributing matches each meta.json, and the skills link entries render
as /open-agent-toolkit/contributing/skills/ and .../docs-tooling/workflows/.
pnpm --filter oat-docs check passes.

* docs(p02-t04): teach the docs skills about Fumadocs nav sync

oat-docs-bootstrap (SKILL.md Sections B and C, AGENTS.md.template) and
oat-docs-authoring's Fumadocs contract now say meta.json is generated by
`oat docs nav sync` from Contents, committed, and strict. oat-docs-apply
changes meta.json only through nav sync and resolves unlisted pages through
approved Contents edits. oat-docs-analyze gains read-only Fumadocs meta.json
checks (missing, stale, rest entries, unlisted pages) and no longer treats
an app without mkdocs.yml as needing migration.

Versions: oat-docs-apply 1.3.1 -> 1.3.2, oat-docs-bootstrap 1.2.1 -> 1.2.2,
oat-docs-authoring 1.0.1 -> 1.0.2, oat-docs-analyze 1.5.2 -> 1.6.0. None
had a bump on this branch versus origin/main; no version pins exist for them.

Failing first: prose-only change with no executable contract; before it,
the skills described nav sync as MkDocs-only and meta.json as optional
hand-authored files, which contradicts p02-t01.
Verify: pnpm oat:validate-skills (65 skills), vitest src/validation (6 files,
391 tests), format:root pass; no tests/ dir exists in the oat-docs-* skills.
Provider views are symlinks; branch `sync --scope project --dry-run` shows
no changes.

* chore(oat): record p02 task ledger before review

* chore(oat): receive p02 review and add p02-t05

* fix(p02-t05): close p02 review findings

M1: add `oat docs nav sync --check`. It computes the navigation, writes
nothing, and exits 1 when a meta.json (or mkdocs.yml) would change or a
page or folder is unlisted, naming each (`stale`, `unlisted` in --json).
apps/oat-docs prebuild runs it after generate-index, so build:docs fails on
stale navigation. Help snapshot and nav sync docs updated.
M2: Fumadocs Contents maps accept .mdx targets, listed by slug; MkDocs
stays .md-only. M3: commands.md no longer implies the hooks write meta.json.
L1: apps/oat-docs/AGENTS.md says meta.json is committed and rewritten only
by nav sync. L2: tests for the documentation.tooling fallback. L3: inline
Markdown is stripped from H1-derived folder titles. oat-docs-analyze
(already bumped on this branch) now points analysis at the --check form.

Failing first: 7 new tests failed before the change (.mdx slug, inline
title, MkDocs check, four --check cases); the help snapshot failed on the
new option. The two L2 tests passed against existing code, so they were
proven by neutralizing: dropping the root match failed the mismatched-root
test, dropping the tooling fallback failed the matching-root test, and
dropping the check write guard failed both writes-nothing tests; all pass
after restore. Negative control: an unlisted docs/zz-probe.md made
`pnpm --filter oat-docs run prebuild` exit 1 naming it; removed, exit 0.
Verify: cli build; vitest src/commands/docs + help-snapshots + release
contract (16 files, 277 tests); nav sync --check on apps/oat-docs exit 0;
oat-docs check; build:docs uncached (0/6 cached) exit 0.

* chore(oat): bookkeeping after p02 pass

* chore(oat): record p02 dispatch and deviations

* chore(oat): record p02 review artifact

* chore(oat): record gate review in project log

* chore(oat): receive p02 gate findings and add p02-t06

* fix(p02-t06): close p02 gate findings

M1: index handling now follows the effective folder metadata. fumadocs-core
16.10.2 gives a root folder (the docs root or a preserved `root: true`) or a
folder with a different `pagesIndex` no implicit landing page, so nav sync
lists `index` first there. Reachability counts a landing page only when the
loader attaches it or `pages` lists it, never just because the file exists.
New fumadocs-loader.test.ts loads the generated meta.json through the docs
app's installed fumadocs-core loader and asserts every page not reported as
unlisted is in the tree, for an ordinary folder and a `root: true` folder,
for both sync and --check, and that --check flags an old root-folder
meta.json the loader would truncate. docs-index-contract.md notes the rule.
M2: oat-docs-analyze grants only Bash(oat docs nav sync --check:*), a
prefix that leaves plain (writing) nav sync ungranted; skills.test.ts pins
the grant, the command form, and the absence of broader oat or Bash grants.
oat-docs-analyze is not an autonomy-contract root, so no prompt site is
re-keyed; it is already bumped (1.6.0) on this branch.
L1: the analysis command is `oat docs nav sync --check --target-dir
<docs-app-dir>`, with the app directory distinguished from its docs/ root.

Failing first: the two root: true loader tests and the grant pin failed
before the change (the ordinary-folder test passed, as expected). Proofs:
ignoring `root: true` in hasImplicitIndex fails both root tests; adding a
Bash(oat docs:*) grant fails the pin; both pass after restore.
Verify: cli build; vitest src/commands/docs + src/validation (21 files,
583 tests); nav sync --check on apps/oat-docs exit 0; oat:validate-skills,
oat-docs check, CLI type-check and check, format:root exit 0.

* chore(oat): bookkeeping after p02 gate pass

* fix(p03-t01): share review-brief source binding across recon helpers

One module (scripts/lib/review-binding.mjs) now owns the review source
projection allowlist, the brief-level source union, and each claim's
source subset. create-review-brief, the artifact-shape contract, and both
validate-packet call sites use it. A verification brief binds when its
sources equal the projected union of its claims' sources and each claim's
evidence and source subset match. Briefs stay blind.

Failing first: on the new two-source fixture (production createReviewBrief
plus reconcileLedger), validatePacket reported 5 REVIEW_BRIEF_MISMATCH
errors (plus one MATERIAL_COVERAGE_ASSURANCE_EXCEEDED, fixed in p03-t02).

Neutralize and restore (review-binding.mjs, once each):
- briefSourcesBind -> return true: the projected-union test fails.
- compare raw manifest sources instead of the projection: the
  allowlist test and the copied-descriptor test fail.
Restored; recon suite 333/333 pass. The per-claim subset check is implied
by the union check plus evidence binding, so neutralizing it alone cannot
fail a test; it stays as the documented per-claim rule.

Bumps recon 1.1.5 -> 1.1.6 with its pins.

* fix(p03-t02): align recon coverage publication with reconciliation

Publication now requires only that every claim named by an accepted
material coverage finding is not verified, which is the downgrade
reconcile-ledger already forces. It no longer also requires the coverage
reviewer's per-statement disposition to be gap, so a reviewer may mark
statements covered while reporting a missing question. packet-contract.md
states the one coverage rule for acceptance, reconciliation, and
publication.

Failing first: on the two-source fixture (coverage marks every statement
covered and reports a material QUESTION_SCOPE_OMISSION), validatePacket
reported MATERIAL_COVERAGE_ASSURANCE_EXCEEDED for the downgraded claim.

Negative control (a verified claim named by a material finding still
fails): integrity-contracts "material coverage gaps prevent verified
assurance". Neutralized the guard (if (false && ...)): that test failed
(333/334 pass). Restored; recon suite 334/334 pass.

* feat(p03-t03): scope recon unresolved issues to claims

unresolvedIssues entries are now a closed union: a string (legacy, read
as global), { text, claimIds } naming a non-empty list of unique claims
the review covers, or { text, scope: 'global' }. Anything else is
rejected at artifact acceptance (INVALID_UNRESOLVED_ISSUE) and is never
read as no issue. reconcile-ledger and validate-packet share
unresolvedIssuesBlockClaim: a scoped issue keeps only its claims below
verified; a global issue keeps every covered claim below verified. The
conditional-routing predicates (any issue present) are unchanged.
packet-contract.md and the worker-contract examples describe the forms.

Failing first: the end-to-end test (production create-review-brief,
reconcileLedger, validatePacket, renderPacket on the two-source fixture)
first threw INVALID_UNRESOLVED_ISSUE while building; with only the shape
accepted, it failed with REVIEW_DISPOSITION_MISMATCH alone (x2, the
affirmed claims). Flipped the two object-rejection tests; added malformed
cases (empty claimIds, non-string ID, uncovered ID, neither, both).

Neutralize and restore (once each; restored suite 338/338 pass):
- validateUnresolvedIssues -> return: both malformed-scope tests fail.
- reconciler drops issueBlocked: scoped-affirmed and global tests fail.
- validator back to any-issue-blocks: end-to-end and scoped tests fail.

* test(p03-t04): prove recon negative controls against helper output

packet-validation's makePacket now builds its verify, adversary, coverage,
and redundant-verify briefs with the production createReviewBrief, so the
per-code negative tests start from helper output. The excerpt controls
use standard packets (which carry those briefs). New controls:
- an edited brief statement, evidence excerpt, locator, or source
  descriptor each fails with REVIEW_BRIEF_MISMATCH and nothing else;
- an edited second-source descriptor in the two-source brief fails;
- a global semantic issue (object or legacy string) left on a verified
  claim fails with REVIEW_DISPOSITION_MISMATCH.
The material-coverage control already starts from helper briefs
(createPacketFixture uses createReviewBrief).

Neutralize and restore (once each; restored suite 343/343 pass):
- reviewBriefBindsClaim -> return true: 8 tests fail (4 brief edits,
  two-source descriptor, union, copied descriptor, claim-bearing briefs).
- excerpt guard (if (false && ...)): wrong-excerpt and paraphrase fail.
- material coverage guard: "material coverage gaps prevent verified
  assurance" fails.
- validator issue guard: the global-issue control fails.

* docs(p03-t05): document Codex agent-limit recovery for recon lanes

recon: a launch rejected before any child is accepted is a
provider/dispatch failure, never a worker failure, and every accepted
artifact is kept. retryLimit now means pre-acceptance admission retries
per lane: with retryLimit >= 1 the controller makes at most one, after
checking completed agents are eligible to be unloaded; 0 allows none.
An alternate route is used only when already approved; otherwise the run
stops partial (PASS_OMITTED lane gap) and asks for a continuation
amendment. No retry or route changes model, effort, role behavior, data
authority, output limits, or reviewer blindness; fresh review lanes stay
fresh; an accepted lane is never rerun to free capacity. Step 5 points to
the dispatch dependency for provider mechanics. The rendered packet
labels the limit "pre-acceptance admission retries per lane"; preview
attempt math is unchanged.

oat-dispatch-subagents provider-codex.md gains the Codex v2 residency
note (eligibility, interrupt_agent vs close_agent, queue-only messages
pinning completed agents, no archive or delete), citing openai/codex
rust-v0.159.2 sources from #333. Appended at the end so the keyed line
in autonomy-contract.md does not move.

skill-contract pins the guidance and that recon does not copy provider
mechanics; replacing the provider-codex pointer fails it (restored).
Bumps oat-dispatch-subagents 1.2.10 -> 1.2.11 with its two pins.

* chore(oat): record p03 task ledger before review

* chore(oat): receive p03 review, record the main merge, add p03-t06

* fix(p03-t06): close p03 review findings

H1: validateReviewBindings now binds a verification brief's claim set in
both directions. Every disposition already had to bind a brief claim; now
every brief claim must also be a distinct ledger claim whose exact
projection binds (statement, evidence, its source subset, and the
brief-level source union). An injected brief claim fails closed with
REVIEW_BRIEF_MISMATCH whether it cites an existing or an injected source.
A real ledger claim the reviewer left without a disposition still binds
(reconciliation keeps it unresolved), so the existing omitted-claim
honest-partial test still passes. Strict disposition equality would have
made that packet unpublishable, so it was not used. Briefs stay blind.
Failing first: both probe shapes from the review (extra claim on
source-1; extra claim plus injected source-3 in the manifest and brief)
validated valid=true. Neutralizing the new check: both probe tests fail.

M2: render-packet adds a Review Downgrades section listing every claim a
review kept below verified (scoped or global issue, coverage finding,
non-affirming disposition) with the review's text, key claim or not. The
fixture's all-key workaround is reverted (key claims: alpha, gamma).
Failing first (renderer at HEAD): 4 tests fail, including a complete
packet with a scoped issue on a non-key claim. Neutralizing
reviewDowngradeLines: 3 tests fail.

L1: the approval preview labels the limit "Pre-acceptance admission
retries per lane (at most one used)" and computes worst-case attempts as
lanes x (min(retryLimit, 1) + 1). packet-contract.md and the recon docs
page state the rule. Failing first (routing at HEAD): the new preview
test fails. Restored suite: 349/349 pass. No version bumps.

* chore(oat): record p03-t06 before re-review

* chore(oat): update autonomous execution learnings

* chore(oat): receive p03 round-2 review and add p03-t07

* fix(p03-t07): close p03 round-2 review findings

H1: the p03-t06 brief binding now covers every brief mode. Each entry
(verification/coverage claims, adversarial provisionalStatements) must be
a distinct ledger claim whose exact projection binds; a repeated claim ID
is its own REVIEW_BRIEF_MISMATCH. New tests: an injected adversarial
controller note, an invented coverage claim, a duplicate adversarial
claim-alpha. Failing first (validator at HEAD): all three published.
Neutralizing the generalization (verify only): those 3 tests fail.

M1: lib/review-omissions.mjs holds one rule for reconciliation and
publication. reconcileLedger (and the reconcile CLI's stdout) now returns
gaps: one material REVIEW_DISPOSITION_OMITTED gap per omitting required
review (semantic, adversarial, coverage), naming the claim and the
review's exact waveId/laneId. Contested and unsupported claims are exempt
(already characterized). validate-packet rejects a packet missing such a
gap (MISSING_REVIEW_OMISSION_GAP); the gap forces partial. Review
Downgrades lists the claim as "not reviewed (no disposition)". Reviewers
may still omit; no accepted review is rerun.
Fixture changes: the two-source fixture records the helper's gaps and
derives status from material gaps. createPacketFixture and the fake run
now have every required review dispose of claim-2 (adversarial
challenged, so it stays contested); before, the production reconciler
silently left claim-2 unresolved and unreviewed in "complete" packets.
Three integrity tests swap a pushed claim-2 disposition for an edit, and
two omission tests use the helper's gaps instead of hand-written ones.
Failing first (validator at HEAD): omitted-disposition and brief-excluded
complete packets published. With the renderer at HEAD, the named-partial
test fails.
Neutralized: validator gap check (2 fail), renderer branch (1 fail),
reconciler gaps (3 fail).

M2: new test, a second claim-alpha with forged evidence on an injected
source-3. Neutralizing only the duplicate-ID clause fails it and the
duplicate adversarial test (2 fail).

Docs: packet-contract.md, SKILL.md step 6, recon docs page. Restored
suite 356/356. No version bumps.

* chore(oat): record p03-t07 before re-review

* chore(oat): record p03 round-3 review and stop at the review cap

* chore(oat): record the operator extension and add p03-t08

* fix(p03-t08): close p03 round-3 review findings

H1: lib/review-binding.mjs owns the request projection (projectReviewScope,
projectReviewQuestions) and each mode's fixed excludedInputs. The generator
uses them, and validateReviewBindings requires reviewBriefRequestBinds for
every brief type (REVIEW_BRIEF_MISMATCH). The schema now fixes brief id (a
lowercase slug, at most 64 chars) and createdAt (UTC ISO-8601).
Brief-field audit (every field a reviewer sees):
- kind, schemaVersion, closed top-level field set: fixed by the schema
- id: lowercase slug <= 64 (new); createdAt: UTC ISO-8601 (new)
- runId: bound to the ledger run; mode: bound to the review kind
- excludedInputs: bound to the mode's fixed declaration (new)
- scope.included/excluded, questions (adversary, coverage): bound to the
  manifest request projection (new); absent from verify (closed schema)
- claims / provisionalStatements: each a distinct, exactly projected ledger
  claim (p03-t06/t07); entry keys closed
- verify evidence (id, sourceId, displayExcerpt, locator): bound to ledger
  evidence; sources: bound to the manifest projection and claim union
Failing first (validator at HEAD): 7 injection tests published
(adversarial and coverage questions, an object question, scope.included,
scope.excluded, verify and adversarial excludedInputs). With contracts at
HEAD the id and createdAt notes published. Neutralizing
reviewBriefRequestBinds: those 7 fail. The packet-contract.md sentence now
lists each field it covers.

M1: the omission exemption is read from the reviews (any rejected, or an
adversarial challenged disposition), never from the published status.
Failing first (omissions module at HEAD): an omitted claim relabeled
contested or unsupported published with no gap. A challenged claim still
needs no gap (control). Neutralizing the exemption: 2 tests fail.

M2: one table-driven test per gap-match clause (code, material, wave, lane,
claim IDs), plus a superset-claimIds control. Neutralizing each clause
alone fails exactly its own case (1 each, 5 runs).

Fixture: createPacketFixture's brief request now mirrors its manifest
request (excludedScope); two-source fixture gains dispositionOverrides,
editReconciled, and editOmissionGaps probe hooks. Suite 375/375.
No version bumps; routing predicates unchanged; briefs stay blind.

* chore(oat): record p03-t08

* chore(oat): record the p03 complexity review and add p03-t09

* refactor(p03-t09): rebuild recon briefs to check integrity and drop omission gaps

Brief integrity is one check. validate-packet rebuilds each brief with the
production createReviewBrief (its id, mode, createdAt, the manifest, the
prior ledger, and the claim IDs the brief lists) and compares canonical
JSON; a throw or a difference is REVIEW_BRIEF_MISMATCH naming the first
differing top-level key. Dispositions must then be unique members of the
brief's claims. Confirmed first: every brief type, contradiction-resolution
included, is created from the prior ledger (fixtures, fake run, SKILL), so
no prior-over-current overlay is needed; priorLedger ?? ledger is used.
Deleted: the duplicate-ID clause, the entry loop, reviewBriefBindsClaim,
briefSourceIds, briefSourcesBind, safeProjectReviewSources,
reviewBriefProjection, reviewBriefRequestBinds. Kept: the projection
allowlists and generator projections, and the id/createdAt schema limits.

Omission-gap protocol deleted (operator decision): review-omissions.mjs,
reconcileLedger's gaps return and manifest parameter, the
MISSING_REVIEW_OMISSION_GAP check, the SKILL.md Step 6 duty, the contract
paragraph, the docs sentence, fixture gap plumbing, and the t07/t08
omission tests. An omitted disposition behaves like uncertain: the claim
stays unresolved, the packet is not forced partial, and Review Downgrades
lists it as not reviewed (small renderer helper). Claim-2 fixture
dispositions stay.

classifyUnresolvedIssue now returns {scope, claimIds} or {error}
(unresolvedIssueDiagnostic folded in; same diagnostics). The reconciler,
validator, and renderer read one affirmingDispositionByReviewKind table.

Tests: one table-driven tamper test (17 rows: statement, excerpt, locator,
descriptor, unprojected descriptor field, injected claim on an existing
and on an injected source, duplicate ID with a forged source, adversarial
note and duplicate, invented coverage claim, questions x2, scope,
excludedInputs, id, createdAt), each REVIEW_BRIEF_MISMATCH. Neutralizing
the rebuild-and-compare check fails all 17 rows. coverage-publication
test deleted (its precondition moved into end-to-end test 1); the two
global-issue tests merged. KD6/KD7 controls and the end-to-end test pass.

Net delta: 15 files, +389 / -1061 (-672); scripts alone +172 / -399.
Suite 357/357. No version bumps; routing predicates unchanged.

* chore(oat): record p03-t09

* chore(oat): record p03 review artifact

* chore(oat): record gate review in project log

* chore(oat): receive p03 gate finding and add p03-t10

* perf(p03-t10): rebuild each recon brief once per validation pass

validateAssurance called checkReviewBrief for every verified claim, so each
brief was regenerated, validated, and hashed N times (gate M1). A per-pass
createBriefChecker now memoizes the rebuild-and-compare result by the exact
brief reference (path and digest) plus review kind; validateReviewBindings
and validateAssurance share it, and it lives only as long as one
compileValidatedRun call (no module-level cache). Every guard is kept:
exact brief reference, mode/run binding, rebuild-and-compare, disposition
uniqueness and membership, and the 17-row tamper table.

compileValidatedRun gains an optional rebuildReviewBrief seam (default:
createReviewBrief) so a test can count rebuilds.
Failing first: the new test counted 12 rebuilds on the two-source packet
(3 bindings + 3 briefs x 3 verified claims) against the expected 3.
Neutralizing the memo (always recompute) fails it again with 12.

Gate scaling probe, validate time at 100/400/800 claims (local):
- before this commit: 102 / 1111 / 4402 ms
- after: 15 / 42 / 97 ms (pre-p03-t09 validator: 29 / 89 / 234 ms)
All packets valid. Suite 358/358. No version bumps.

* chore(oat): bookkeeping after p03 gate pass

* fix(p04-t01): recompute next's exit-gate fingerprint with v2 exclusions

oat-project-next 5.0 recomputed every stored fingerprint with only the
state-carrier exclusion, so a fresh effective-delta-v2 generation read as
stale after any .oat/projects or .oat/repo commit. It now recomputes v2 with
the implement skill's three literal exclusions and v1 with v1 rules, and
points at completion-and-closeout.md Step 14 where the algorithm lives.

A contract test extracts the v2 pathspec sentence from both skills and
asserts the sets are identical, and pins next's v1 rule to the carrier only.
Bump oat-project-next 1.1.3 -> 1.1.4 and its pins.

Failing-first: new test failed before the edit ("Expected one
effective-delta-v2 exclusion pathspec sentence, found 0").
Neutralize-and-restore: dropping .oat/repo from next's v2 sentence fails
the test (2 vs 3 pathspecs); swapping next's v1 pathspec for .oat/projects
fails it too; both restored, test passes.

Verification: vitest src/commands/init/tools/shared src/validation
(isolated HOME) 902/902 pass; check:skill-bumps exit 0.

* feat(p04-t02): add oat project closeout-check and guard complete-state

New read-only `oat project closeout-check <path> [--autonomous] [--json]`
evaluates the closeout invariant from state.md on disk. A snapshot is required
when the effective layered workflow.postImplementSequence is set, when the
run is autonomous (--autonomous or OAT_AUTONOMOUS=1 via an injectable env),
or when oat_workflow_mode is lite. Once a snapshot exists its record is
authoritative and config is never consulted. It reports status, the missing
invariant, route oat-project-implement, and the next owner (stored step and
skill, or approval); incomplete exits 1. complete-state takes --autonomous,
runs the same check before any write, and refuses with the same message.
Legacy projects recover through oat-project-implement; no override flag.

Failing-first: complete-state refusal tests (cases 1 configured/flag/lite,
2, 3, 6, 7 and --autonomous) failed 8/16 before the guard existed.
Neutralize-and-restore, one clause each (all restored, 49/49 green):
- env read -> false: case 7 fails in closeout-check and complete-state
- complete-state guard skipped: cases 1, 2, 3, 6, 7 and the trace fail
- configured input ignored: configured cases and the trace fail
- lite input ignored: lite cases fail
- stored order -> canonical order: disk transition trace fails
- approval transition removed: case 3 and the trace fail
- terminal status rule removed: case 4c and the trace fail
- malformed snapshot -> complete: all 8 malformed variants fail
- config consulted with a snapshot: snapshot-authority tests fail

Transition trace: configured and snapshot-absent state on disk, stored
order document, summary, pr; a fresh invocation after every Step 15 write;
interruption (summary.md in_progress) and a failed boundary both resume at
summary; complete-state refuses until status complete, then succeeds.
Probe (branch build): recon-rework complete; claude-effort-levels malformed
(hand-written status: pending, approval: null); this project snapshot_missing.

Verification: cli build exit 0; vitest src/commands/project and
help-snapshots (isolated HOME) 1313/1313; oat-docs check exit 0; cli check
and type-check exit 0.

* feat(p04-t03): route terminal closeout through the closeout check

Every terminal consumer now runs `oat project closeout-check` (adding
--autonomous under OAT_AUTONOMOUS=1) and routes an incomplete closeout to
oat-project-implement with the reported invariant and next owner:
- implement Step 15 commits the snapshot and checks it before the first
  sequence child; a missing or malformed snapshot dispatches nothing.
- implement Step 16 runs the check before any completion write and
  continues only on complete or not_required.
- next 5.1 runs the check before every later route, so a configured,
  autonomous, or lite closeout with no snapshot now routes to implement.
- complete adds Step 1.5 before its upfront questions and any write, and
  re-checks before complete-state, which now gets --autonomous too.
oat-project-autonomous and its gate inventory are unchanged: it completes
only through implement and complete. The control-plane recommender and the
state dashboard stay out of scope and rely on complete-state's refusal.
Bump oat-project-complete 1.7.13 -> 1.7.14 and its pins; implement and next
are already bumped on this branch. Named-skill matrix rows reclassified.

Failing-first: the four new contract tests failed before the edits
(missing markers in implement Steps 15/16, next 5.1, complete Step 1.5).
Neutralize-and-restore, one consumer each (all restored, green):
- implement Step 15 pre-dispatch check removed: its test fails
- implement Step 16 check removed: its test fails
- next 5.1 check removed: its test fails
- complete Step 1.5 gate removed: its test fails
- complete Step 5 re-check removed: its test fails

Verification: vitest src/commands/init/tools/shared src/validation
(isolated HOME) 906/906; node --test oat-project-implement 37/37;
oat:validate-skills exit 0; check:skill-bumps exit 0.

* feat(p04-t04): add operator-only exit-gate waivers

DR-260927-operator-waiver-for-test-only: no automatic test-only exception.
oat_implement_exit_gate gains an append-only `waivers` list; each entry
records waived_by, reason, from_commit..to_commit (Git range), the
covered_fingerprint at to_commit in the generation's own version, and a
UTC waived_at. A waiver never rewrites reviewed_head, the implementation
or freshness fingerprints, freshness_head, or an earlier waiver, and never
revives a generation already persisted stale. It is written only on an
explicit operator instruction, never inferred or self-issued; under
OAT_AUTONOMOUS=1 every waiver write is refused and the run stops. A waived
generation reads allowed only while nothing substantive lands after the
covered range; v1 and v2 share one rule and differ only in which paths are
substantive. Malformed or unverifiable waivers fail closed as stale.
next 5.0 validates waivers read-only; summary and pr-final list every
waiver; the state template and three docs pages describe the field.
Bump oat-project-pr-final 1.6.6 -> 1.6.7 and pins; summary, implement,
and next are already bumped on this branch.

Failing-first: the six new waiver tests failed before the edit (missing
Operator waivers section).
Neutralize-and-restore, one clause each (all restored, green):
- operator-instruction clause removed: issuance test fails
- OAT_AUTONOMOUS refusal removed: issuance test fails
- append-only clause removed: record test fails
- malformed fail-closed weakened to "ignored": v1/v2 rule test fails
- later-substantive-stale clause removed: v1/v2 rule test fails
- per-version covered fingerprint removed: v1/v2 rule test fails
- pr-final listing made optional: summary/PR test fails
- reason dropped from the skill record: record test and the v1 and v2
  executable checks fail (the check reads its field set from the skill)
- executable check without waiver validation: v1 and v2 checks fail
- executable check covering past to_commit: v1 and v2 checks fail

Verification: vitest src/commands/init/tools/shared, src/commands/project,
src/validation (isolated HOME) 2162/2162; oat:validate-skills exit 0;
check:skill-bumps exit 0; oat-docs check exit 0.

* chore(oat): record p04 task ledger before review

* chore(pjm): run a complexity review when review or gate budgets run out

Adds BL-261001-run-a-complexity-review-when as the first slice and gives BL-260818 the
complexity-review input and a fourth disposition, simplify.

* chore(oat): receive p04 review and add p04-t05

* fix(p04-t05): close p04 review findings

Review: reviews/archived/p04-review-2026-10-01T174412Z.md (M1, L1-L3).

M1: the waiver is now reachable. Operator waivers gain a stale-boundary
offer: before an allowed qualified generation that is stale only through
uncovered descendants is persisted stale, an interactive run lists the
commits and asks the operator to waive that range or start a new gate run;
stale is persisted only when the operator declines. Under OAT_AUTONOMOUS=1
no waiver is offered or issued: persist stale and start a new gate run.
Other stale causes get no offer. An operator may also record a waiver
before resuming implement. Step 14 item 1 and the Step 15 freshness
re-check route through the offer; next 5.0 mirrors it in its stale note.
Docs (implementation-execution, workflow-gates) describe both paths.
L1: transition-trace step 3 wrote summary.md, which the check never reads;
removed. The failed boundary is now the interruption step.
L2: the BL-260829 comment moved back above its own describe.
L3: an approval_pending nextOwner now carries both writes, `approval:
approved` and `approval: not_required`; the message names both; Step 15
says nextOwner is a step to dispatch or, with an empty pre_approval, the
approval boundary to record. Autonomy coverage maps the three new L3
lines to NG; named-skill matrix rows added for the two new sentences.

Failing-first: 3 new tests failed before the edits (offer clauses, Step 15
nextOwner wording, case 3c empty pre_approval).
Neutralize-and-restore (each restored, green):
- interactive offer removed: offer test fails
- autonomous never-offer clause removed: offer test fails
- Step 14 item 1 bypasses the offer: offer test fails
- next mirror removed: offer test fails
- approval writes dropped from nextOwner: case 3c fails

Verification: cli build exit 0; vitest src/commands/project,
src/commands/init/tools/shared, src/validation (isolated HOME) 2165/2165;
check:skill-bumps exit 0; cli check, type-check, validate-skills exit 0.
No version bumps.

* chore(oat): bookkeeping after p04 pass

* chore(oat): record p04 review artifact

* chore(oat): record gate review in project log

* chore(oat): receive p04 gate findings and add p04-t06

* fix(p04-t06): close p04 gate findings

Gate: reviews/archived/p04-review-2026-10-01T180727Z.md (M1, L1).

M1: a failed snapshot chose `prePending ?? postPending` before checking
approval, so with pre-approval work done, approval pending, and retro
stored post-approval it named retro. The failed branch now picks its owner
in approval-aware order: pending pre-approval work, the pending approval
boundary (both writes named), post-approval work, then status repair. A
post-approval step is never named while approval is pending. The
approval owner is shared with the approval_pending branch.
L1: implementation-execution and lifecycle docs now say an autonomous run
that finds a stale generation persists stale and starts a new gate run;
autonomy refuses only an attempted waiver write.

Failing-first: the gate's exact fixture, added as a closeout-check
command-boundary regression, failed before the fix (nextOwner was the
post_approval retro step). Two ordering companions (pre-approval work
first; post-approval work once approval is recorded) pin the other legs.
Neutralize-and-restore: putting post-approval work ahead of the approval
boundary again fails the fixture test; restored, 32/32 green.
Branch-CLI probe with the gate's fixture and command: sequence_failed,
nextOwner kind approval with both writes, exit 1.

Verification: cli build exit 0; vitest src/commands/project (isolated
HOME) 1254/1254; oat-docs check exit 0; cli check and type-check exit 0.
No version bumps.

* chore(oat): bookkeeping after p04 gate pass

* fix(p05-t01): never overwrite a CLAUDE.md that AGENTS.md links to

Under a shim strategy with --force, instructions sync overwrote a CLAUDE.md
that an AGENTS.md resolves to, destroying the only copy of the instructions.
Planning now keeps such a file and reports it as a skip. The resolves-to
check (hard link, or a symlink chain ending at CLAUDE.md) is extracted from
inspectManagedShim and shared with the new findAgentsResolvingTo helper; a
CLAUDE.md that is itself a symlink counts only when an AGENTS.md chain passes
through it, so the ordinary CLAUDE.md -> AGENTS.md shim is still replaced.
Single planning-time guard; no apply-time re-check.

Failing-first: 5 new integration tests failed before the guard (pointer and
symlink --force on AGENTS.md -> CLAUDE.md, a symlink chain, a hard link, and
a cross-directory link); the non-linked control passed before and after.
Neutralize-and-restore: making keepClaudeFilesAgentsResolveTo return its
input unchanged failed the same 5 tests (143 passed); restored, 148/148 pass.

Verify: HOME=$(mktemp -d) vitest run src/commands/instructions -> exit 0.

* refactor(p05-t02): make dispatch record validate-only

Per DR-260927-dispatch-record-validates, remove dispatch-record persistence and keep the
validate-only command plus its schema modules for the managed Claude validation path.

- record.ts: delete the <project>/dispatch/ journal writer, writer lock, revision naming,
  fallback-claim publication, trigger and related-record reads; recordProjectDispatch now
  takes { input } and returns { status: 'validated-only', record, runtimeIdentity }.
- index.ts: drop --project, the persisted status, and the <project> redaction label.
- record.test.ts: prune journal tests (61 -> 49 cases); move the journal-byte redaction
  assertions (absolute, assignment-form, nested evidence, rollout, Claude answer) onto
  the command's validate-only JSON output via validatedOutput(); add a fallback-link
  refusal case.
- Skills and docs: remove persistence wording from oat-dispatch-subagents (SKILL.md,
  record-schema.md, managed example), oat-project-dispatch-subagents, review-provide,
  review-provide-remote, plan-writing, implement dispatch-and-dry-run.md, CLI reference,
  evidence-layers, orchestration-model (Journal participant -> run record),
  implementation-execution, scope-and-surface.
- Bumps: oat-project-dispatch-subagents 1.1.7, oat-project-review-provide 1.5.12,
  oat-project-review-provide-remote 1.1.9, oat-project-plan-writing 1.2.35; pins updated.

Failing-first: the rewritten skills.test.ts case (persistence wording absent) failed on
oat-dispatch-subagents/SKILL.md before the doc edits.
Neutralize-and-restore: appending the old opt-in sentence to plan-writing SKILL.md failed
that case; re-adding a --project option failed the help-snapshot case; both restored.
Probe: dist project dispatch record --project x --event-file - -> unknown option, exit 1.
Search from the plan: no output, rg exit 1.
Verify: build 0; vitest dispatch+help+validation 710/710; test:smoke 163/163;
test:skills 689/689; oat:validate-skills 0.

* feat(p05-t03): let test-only package changes skip the lockstep bump

Per DR-260927-test-only-paths-skip, each public package's versionPolicyIgnorePatterns now
lists exactly the test paths its tsconfig.json excludes from dist: cli src/**/*.test.ts and
src/**/__tests__/** (beside assets/**), control-plane **/*.test.ts and **/*.spec.ts,
docs-config and docs-transforms src/**/*.test.ts, docs-theme none. AGENTS.md Package
Management states the rule.

Tests: a contract test reads each package's tsconfig exclude (minus node_modules, dist) and
requires it to equal the ignore patterns; path cases for every package; release-utils and
runVersionBumpCheck cases for a cli test-only change (no bump) and a non-test src change
(bump required, the negative control).

Failing-first: 4 new cases failed before the pattern change (contract equality, path
cases, test-only diff, test-only check). Neutralize-and-restore: widening the cli pattern to
src/** failed 5 cases including the negative control; restored, 77/77 pass.

Real probe in a scratch worktree off origin/main (contract change applied uncommitted):
test-only commit to packages/cli/src/release/release-utils.test.ts ->
release:check-versions exit 0 ("no public package changes"); then a commit to
packages/cli/src/index.ts -> exit 1 ("Changed packages: @open-agent-toolkit/cli").
Worktree removed with git worktree remove --force. Deviation: the probe ran
pnpm install --frozen-lockfile instead of worktree:init; see the phase report.

* feat(p05-t04): report YAML errors with location and check key types

parseSkillFrontmatter now uses a LineCounter and returns the first parser error's position
(or the non-mapping metadata node's position) as a problem. skill-frontmatter-unreadable
appends it as a SKILL.md line and column, e.g. "...: line 3, column 14: Nested mappings
are not allowed in compact mappings". New findSkillFrontmatterTypeProblems checks parsed
types of name, description, allowed-tools (string) and disable-model-invocation,
user-invocable (boolean); the oat-* loop reports skill-frontmatter-type findings beside the
existing semantic checks, which are unchanged.

Failing-first: 4 new cases failed before the change (bare-colon description, name: 123,
user-invocable: "true", metadata: [1.0.0]); the valid-fixture control passed. Five
existing malformed-frontmatter expectations gained the location suffix.
Neutralize-and-restore: dropping the problem from the finding and skipping type problems
failed the same 4 cases; restored, 622/622 pass.

Verify: HOME=$(mktemp -d) vitest src/validation src/commands/shared -> 622 pass;
pnpm oat:validate-skills -> 65 skills OK.

* fix(p05-t05): route quick-mode discovery to quick-start

The status/list recommender and the state dashboard routed quick-mode discovery through
oat-project-plan, a two-hop route that disagreed with the oat-project-next and
oat-project-progress skill tables. Changed exactly: router.ts discovery:in_progress:2 and
discovery:complete:1, and generate.ts quick:discovery:complete, to oat-project-quick-start.
Router discovery:in_progress:3 and the dashboard's quick:discovery:in_progress stay at
oat-project-discover, now pinned by tests. Quick plan:in_progress routing is untouched
(BL-261001-route-quick-mode-plan).

Mechanical propagation (outside the declared files): split/__tests__/run.test.ts pinned the
old route for a seeded quick child's project status; it now expects quick-start.

Failing-first: 3 router cases and 1 dashboard case failed after the pins were flipped and
before the route change. Neutralize-and-restore: pointing router discovery:in_progress:3 at
quick-start failed "keeps quick discovery tier 3 in discovery"; pointing the dashboard's
quick:discovery:in_progress at quick-start failed the dashboard case; both restored.

Verify: control-plane vitest src/recommender 49/49; HOME=$(mktemp -d) cli vitest
src/commands/state 31/31; src/commands/project/split 39/39.

* fix(p05-t02): prune dispatch-record persistence tests missed in p05-t02

Recovery event bw3-p05-recovery-1 (attempt 1/10, phase-standing) for original commit
a92f397fd. p05-t02 removed --project and the persisted status, but three test files outside
its verification set still exercised persistence and failed:
commands.integration.test.ts (2 cases), e2e/workflow.test.ts (2 cases), and
init/tools/shared/review-skill-contracts.test.ts (pinned the removed opt-in wording).
Bounded, test-only correction: run the command without --project, expect
status validated-only, assert on stdout instead of journal bytes, and assert the
persistence wording is absent. No production code changes.

Verify (pre-commit): focused vitest src/commands/project src/e2e
src/commands/commands.integration.test.ts 1362/1362; phase vitest src/commands
src/validation src/release 5979/5979; whole CLI suite 8012/8012.
Ledger: state.md p05 pending_attempt marked completed.

* docs(p05-t06): narrow the packs inventory redaction claim

troubleshooting.md claimed "Reported project and home paths remain redacted" for the
packs:inventory diagnostic. The code redacts less, so the entry now says what it does:
the project root (to a relative path or .) and the home root (to ~) are replaced only
when that scope is part of the run, by literal replacement of the exact root path, and a
path outside those roots, such as a global bundle path in an assets error, stays absolute.

Code paths (packages/cli/src/commands):
- status/index.ts unavailablePackReport (~204-240): replaceAll over projectRoot/userRoot
  pairs built only from the run's scopeRoots; collectPackReport (~708-713) builds roots.
- tools/shared/format-pack-inventory.ts redactPackText (~23-62): same literal replacement
  for unavailablePackEvidence reasons.
- doctor/index.ts (~1098-1103 roots, ~1186-1213 packs:inventory check): message is the
  evidence diagnostic detail from redactPackText.

BL-260903-verify-the-packs-inventory: placeholder acceptance criteria filled with this
outcome. No new test (docs-only, per plan). Verify: pnpm --filter oat-docs check -> 0.

* chore(oat): record p05 task ledger before review

* chore(oat): receive p05 review and add p05-t07

* chore(pjm): file wave 3 follow-ups for bundle-assets and oat-wrap-up

* chore(oat): fix the bundle-assets backlog reference

* fix(p05-t07): close p05 review findings

M1 (tools/release/release-utils.ts): findChangedWorkspaceDirsFromPaths judged a path under
a dependency root by the dependent's ignore patterns, so a control-plane or docs-transforms
test-only change still marked packages/cli or packages/docs-config. Each path inside a
public package is now judged by the patterns of the package that owns it; paths outside
every package (additional roots) keep the contract's own patterns.
- Failing-first: with the real dependency map, the control-plane and docs-transforms
  test-only cases failed; the non-test negative control (router.ts -> cli + control-plane;
  docs-transforms index.ts -> docs-config + docs-transforms) passed before and after.
- Neutralize-and-restore: dropping the owner lookup failed the 2 test-only cases; restored,
  src/release 80/80.

M2: delete journal-only machinery with no production caller.
- fs/io.ts: withContainedWriterLock, journalErrorCode, redactedFsError,
  publishContainedJsonRevision and their private helpers; io.test.ts lock/publication cases
  and the journal race mocks.
- absolute-paths.ts: assertJournalIdentityHasNoAbsolutePath; journal docstring reworded
  (also two journal comments in generic-dispatch-record.ts).
- oat-dispatch-record.ts: triggerRecord/relatedRecords inputs, the fallback-claim and
  fallback-link event kinds and branches, and the fallback control-field lists and digest
  only they used, with their tests. The record keeps fallbackClaim: null and a
  not-applicable fallback, so the validated output shape is unchanged.
- Failing-first: record.test.ts now expects fallback-link and fallback-claim to be refused
  as unknown kinds (Invalid discriminator value); it failed before (got "Fallback requires
  the rejected trigger record"). The review's rg for the deleted names went from 60 matches
  to none.
- Unchanged validate-only/managed path: branch CLI on managed-claude-example.json
  implementer and reviewer inputs -> exit 0, status validated-only, keys
  record/runtimeIdentity/status, payload.variant intact; managed and dispatch tests pass.

Verify: cli build 0; vitest src/release src/fs src/providers src/commands/project/dispatch
1102/1102; test:smoke 163/163; cli type-check 0; pnpm lint 0 (0 cached); whole CLI suite
7954/7954; pnpm check 0. No skill text changed, so no version bumps.

* chore(oat): bookkeeping after p05 pass

* chore(oat): record p05 review artifact

* chore(oat): record gate review in project log

* chore(oat): receive p05 gate finding and add p05-t08

* fix(p05-t08): compare inodes through symlinked AGENTS.md

Gate M1: the --force preservation check compared realpaths when AGENTS.md was a symlink, so
pkg/AGENTS.md -> ../alias.md, with alias.md hard-linked to CLAUDE.md, was missed and --force
replaced the protected CLAUDE.md. agentsResolvesToClaude now resolves a symlinked AGENTS.md
to its endpoint and compares that endpoint's device and inode with the regular CLAUDE.md.
Read failures still fail closed (unable to resolve AGENTS.md; kept). Callers pass only a
regular CLAUDE.md, so the ordinary CLAUDE.md -> AGENTS.md shim is still overwritten.
The unused claudePath parameter is dropped.

Failing-first: the new command-boundary regression (root AGENTS.md independent, alias.md a
hard link of CLAUDE.md, pkg/AGENTS.md -> ../alias.md, sync --force) failed for pointer,
symlink and copy; each now skips the CLAUDE.md action, and its content and inode are
unchanged. Direct-symlink, chain, hard-link and ordinary-overwrite controls still pass.
Neutralize-and-restore: making the endpoint comparison always false failed the 3 new cases
and the 5 symlink-based keep cases (8 failed); restored, 151/151 pass.

Verify: HOME=$(mktemp -d) vitest run src/commands/instructions -> exit 0, 151/151;
cli check and type-check 0.

* chore(oat): bookkeeping after p05 gate pass

* chore(p06-t01): bump lockstep public packages to 0.3.11

Follows the version files #334 (98d1d5246) changed: the five public package.json
files and packages/cli/assets/public-package-versions.json. The sync manifest's
oatVersion is a sync stamp, not a lockstep version, and is left unchanged.

Verified: release:check-versions exit 0 against origin/main 98d1d5246;
release:validate exit 0 (5 packages at 0.3.11); scoped release and bundle tests pass.

* chore(p06-t02): archive the backlog items shipped in wave 3

Archived 13 items with the branch CLI (oat backlog archive, all exit 0). BL-260909 cites
DR-260927-dispatch-record-validates; BL-261001-recover-recon-lanes-after cites discovery
Key Decision 8; BL-260829 cites the p01/p02 Step 7a reviews (de9c98848, c129e82ab) under
oat-project-implement 2.3.14 (origin/main 8f6d5b1d2). BL-260806 stays open until closeout.

The BL-260829 archive rewrote two inbound references. Adds the Wave 3 curated overview
note; regenerate-index is idempotent; pjm doctor reports only its pre-existing warning.

* chore(oat): record p06 task ledger before review

* chore(oat): close p06 review findings and write the PR hand-off

* chore(oat): record p06 review artifact

* chore(oat): record gate review in project log

* chore(oat): prepare final implementation closeout

* chore(oat): record the final review and close its findings

* chore(oat): persist exit gate generation 1 intent

* chore(oat): record exit gate generation 1 acceptance

* chore(oat): record final review artifact

* chore(oat): record gate review in project log

* chore(oat): persist exit gate generation 1 result and receive intent

* chore(oat): receive exit gate generation 1 review

* chore(oat): exit gate generation 1 allowed

* chore(oat): persist the post-implement sequence snapshot

* docs: generate summary for backlog-wave-3

Promote five key decisions to reference/decisions (DR-261001-*); roll up 20 project-log entries.

* chore(oat): record the summary closeout step

* chore(pjm): record backlog wave 3 in current state and roadmap

Add the CLI 0.3.11 entry to current-state.md and drop the closed
BL-260829-order-phase-bookkeeping-before bullet from the roadmap's Next list.

* docs(backlog-wave-3): update documentation from project artifacts

- contributing/code.md: test-only changes skip the lockstep bump (replaces the stale claim)
- troubleshooting: older oat rejects template resolve and closeout-check; need 0.3.11+
- lifecycle: oat-project-complete runs closeout-check before any completion write
- instruction-sync: --force keeps a CLAUDE.md that an AGENTS.md resolves to
- recon: brief rebuild-and-compare, scoped unresolvedIssues, coverage gaps downgrade
- add-docs-to-a-repo: the Fumadocs scaffold prebuild does not run nav sync --check

* chore(backlog-wave-3): mark docs updated

* chore(oat): record the document closeout step

* chore(oat): archive residual plan reviews and mark the final PR ready

Archive the two fixes_completed plan artifact-review files under reviews/archived/ and
rewrite their plan.md ledger and implementation.md references (pr-final Step 0.5).
Set oat_pr_status to ready now that the final PR artifact exists.

* chore(oat): record final PR metadata

Final PR opened as #336 (https://github.com/voxmedia/open-agent-toolkit/pull/336).
Set oat_pr_status to open, oat_pr_url, and oat_phase_status to pr_open, and route the
next milestone to oat-project-revise or either oat-project-complete ordering.

* chore(oat): record the pr closeout step

* docs(oat): build backlog-wave-3 project recap

Unattended project-recap run e33c16db-fd14-47c7-9c4e-2cdccfef6d9b, outcome built.
Host verify rung: headless Chrome at 320, 768, and 1440 pixels; all browser-free checks pass.

* docs(oat): record the recap outcome in the backlog-wave-3 summary

* docs(oat): rebuild backlog-wave-3 project recap

Rebuild after d4fe765f8 added the Explainer Outcome section to summary.md, an approved input.
Unattended project-recap run 810e34a7-ca5d-4437-a88a-e1cd02ec37a7 replaces e33c16db in place.
Outcome built; host verify rung at 320, 768, and 1440 pixels; all browser-free checks pass.

* chore(oat): closeout awaiting final approval

* chore(oat): record autonomous final approval

* chore(oat): closeout sequence complete

* chore(oat): mark implementation complete

* chore(oat): record the BL-260806 live closeout trace

* chore(pjm): archive BL-260806 with the wave 3 closeout trace

* docs(oat): rebuild backlog-wave-3 project recap for completion

Rebuild after implementation.md gained the final-approval record and the BL-260806 closeout trace.
Unattended project-recap run b2bc7235-9374-4a84-84a4-3abbc6b2ebe2 replaces 810e34a7 in place.
Outcome built; host verify rung at 320, 768, and 1440 pixels; all browser-free checks pass.

* chore(oat): complete project lifecycle for backlog-wave-3
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant