Skip to content

test: qualify native domain migration in linked worktrees - #128

Open
roodboi wants to merge 4 commits into
nextfrom
codex/native-domain-acceptance
Open

roodboi wants to merge 4 commits into
nextfrom
codex/native-domain-acceptance

Conversation

@roodboi

@roodboi roodboi commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

The native frontend acceptance driver previously checked one linked-worktree hostname. Add an opt-in --domain-migration sequence using real Doctor commands and inferred branch identity to qualify legacy/OAuth/service aliases, retaining restarts and exact rollback. It requires unchanged run/owner/plan/volume bindings, retained marker, exact original file bytes/modes, an unchanged primary checkout, drift refusal, and removed-route rejection bracketed by a live positive control.

The harness preserves failed fixtures. Only audited read-only status/inspect commands may repeat an exact, bounded structured provider_busy refusal within a fixed budget; exec, refresh, mutations, malformed responses and timeouts never replay. Overall success is written only after owned disposal succeeds. Production runtime behavior is unchanged.

Validation: independent harness review found no actionable blocker; 24 Python controls, Python compilation, privacy and diff checks pass; all eight CI jobs pass on 4548f6a84edc5e19cde6a5fd471429763d3cfcae. Controls cover missing OAuth routes, replaced volumes, rollback drift/overwrite, wrong restoration, transport failures and bounded observation behavior.

Native qualification: the complete Apple-silicon macOS matrix passed 54 behavior checks and owned disposal using this harness at 4548f6a8, private frontend composite d67f4a9d (authority admission correction plus release diagnostics and timeout-test correction), and native source e3f3d07b. This includes prepared-base linked startup, source edit/down/up, exact retained data, legacy/OAuth/service aliases, forward migration/restart, drift refusal, exact rollback/reverse restart, removed-route rejection, and primary checkout byte/mode preservation. The pool, base, alias and fixture data were removed; post-run inspection found zero fixture processes and a free HTTPS port. Earlier runtime failures remain retained evidence; this test-only PR does not supply the production fixes.

Release signal: test infrastructure only. Explicit private-CA loopback HTTPS does not prove system DNS, browser trust, concurrent sibling readiness, comparative performance, or complete HACK-1163 acceptance. No DNS/trust or publishing changes. Tracks HACK-1163, which remains In Progress. Base: next.

@linear-code

linear-code Bot commented Oct 4, 2026

Copy link
Copy Markdown

HACK-1163

@roodboi

roodboi commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Native qualification checkpoint:

  • The first attempt passed prepared-base startup, exact linked source sharing and HTTPS, then failed during dependency refresh before the marker exec was dispatched. The frontend had discarded the native error identifiers. Its original cause remains unproven; the uncertain fixture is preserved. fix: retain safe native dependency refresh diagnostics #129 restores only reviewed diagnostic codes without replaying requests.
  • A fresh supervised run with the diagnostic frontend passed 23 checks, including marker exec, source editing, migration preview/application and the forward retaining restart. The next read-only graph inspect returned structured provider_busy; a later observation confirmed ready-observed with all six expected aliases. The supervised foreground remained alive. Normal owned cleanup and disposal then completed, with receipts retained privately.
  • Commit 833c174d adds a 15-second, evidence-preserving re-observation window solely for exact graph inspect and runtime status contention. Other errors, malformed replies, timeouts, mutations, exec and refresh are not retried. All 24 Python controls pass. The complete native forward/reverse run with this correction is still pending; browser/DNS trust and simultaneous sibling-runtime acceptance remain separate gates.

@roodboi

roodboi commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

The next supervised full run used harness 833c174d and the PR #129 diagnostic frontend. The first startup, retained marker and ordinary shutdown passed, but the second startup after the host edit failed before the domain migration sequence: the graph became ready, while the new shared HTTPS owner never published its endpoint.

The previous owner had retired cleanly. The visible missing-endpoint exception came from cleanup and masked the new owner's acquisition failure; the underlying helper-startup cause remains unproven. Both failed pools are now stopped with disks, ownership records and private evidence preserved. No failed mutation was replayed.

PR #130 preserves a bounded, generation-bound startup diagnostic across later cleanup failures. It changes diagnosis only; it does not treat missing endpoints as recovery authority or claim to fix the underlying failure. The complete linked-worktree domain round trip remains an open native gate.

@roodboi

roodboi commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Native qualification checkpoint (M3, October 4): harness 833c174d against a private composite of PR #129 (ac36203f) and PR #130 (1e144866).

22 native assertions passed: prepared linked-worktree startup, edit/down/up with retained marker, all four initial legacy/OAuth/service aliases, migration preview/apply, primary-checkout isolation, and unchanged live graph bindings before explicit restart. The forward restart then refused replacement startup because the previous HTTPS lease finalization remained unconfirmed. Full forward/reverse domain acceptance has not passed; keep this PR draft.

The stopped graph has clean journal, absent container/network, retained data and confirmed relay cleanup. Captured owner/run/boot/control identities and exposed release predicates align; they do not identify the transient release refusal. The owner currently destroys the socket on that error, discarding its cause. Provider contention is only a hypothesis. Next correction is bounded, lease-bound release diagnostics with unchanged cleanup authority and no automatic replay.

The owned VM was stopped normally and zero fixture processes remain. Disks, receipts and private evidence were retained. No Event Agent or stable resources changed. PR #130 has all eight CI jobs passing, but this does not establish a native startup/cleanup fix.

@roodboi
roodboi marked this pull request as ready for review October 5, 2026 01:29

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant