Issue Description
Propose replacing claude-sonnet-4.6 with claude-sonnet-5 in the baseline-equivalence model list while retaining gpt-5.6-luna and the existing validation gates.
During investigation of PR #2894, Sonnet 4.6 was absent from the model catalog returned under the CI configuration, and session creation failed with model-unavailable. A subsequent isolated test targeting Sonnet 5 passed across all four tested Vally/Copilot SDK combinations, including the repository’s current versions.
Proposed Change
Update Resolve-ModelList in scripts/evals/Invoke-BaselineEquivalence.ps1 so the calibration and ci tiers use:
gpt-5.6-luna
claude-sonnet-5
Update associated model-selection tests, help text, and current documentation. Preserve historical results and references that describe earlier runs.
Keep the comparison judge, retry policy, execution-health checks, evidence-completeness checks, invariants, and tier-specific gating behavior unchanged. No Vally or Copilot SDK dependency change is proposed.
This retains cross-model coverage rather than reducing the suite to Luna-only.
Additional Context
The isolated harness changed only the requested model ID. It used the existing CI credential configuration, a fixed arithmetic prompt, no permitted tools, and the same CLI version.
| Vally |
Copilot SDK |
Sonnet 4.6 |
Sonnet 5 |
| 0.14.0 |
1.0.11 |
Model unavailable |
Pass |
| 0.15.0 |
1.0.9 (current) |
Model unavailable |
Pass |
| 0.14.0 |
1.0.9 |
Model unavailable |
Pass |
| 0.15.0 |
1.0.11 |
Model unavailable |
Pass |
All Sonnet 5 cells reported catalog presence, the expected response, zero tool-permission requests, and successful cleanup. Recorded per-cell dependency-lock and CLI fingerprints matched the earlier Sonnet 4.6 test.
Evidence:
The smoke establishes basic availability, not full RPI behavior or calibration success. It does not prove why Sonnet 4.6 became unavailable. A separate Luna nonzero exit remains unresolved and should not be considered fixed by this change.
Acceptance Criteria
Issue Description
Propose replacing
claude-sonnet-4.6withclaude-sonnet-5in the baseline-equivalence model list while retaininggpt-5.6-lunaand the existing validation gates.During investigation of PR #2894, Sonnet 4.6 was absent from the model catalog returned under the CI configuration, and session creation failed with
model-unavailable. A subsequent isolated test targeting Sonnet 5 passed across all four tested Vally/Copilot SDK combinations, including the repository’s current versions.Proposed Change
Update
Resolve-ModelListinscripts/evals/Invoke-BaselineEquivalence.ps1so thecalibrationandcitiers use:gpt-5.6-lunaclaude-sonnet-5Update associated model-selection tests, help text, and current documentation. Preserve historical results and references that describe earlier runs.
Keep the comparison judge, retry policy, execution-health checks, evidence-completeness checks, invariants, and tier-specific gating behavior unchanged. No Vally or Copilot SDK dependency change is proposed.
This retains cross-model coverage rather than reducing the suite to Luna-only.
Additional Context
The isolated harness changed only the requested model ID. It used the existing CI credential configuration, a fixed arithmetic prompt, no permitted tools, and the same CLI version.
All Sonnet 5 cells reported catalog presence, the expected response, zero tool-permission requests, and successful cleanup. Recorded per-cell dependency-lock and CLI fingerprints matched the earlier Sonnet 4.6 test.
Evidence:
The smoke establishes basic availability, not full RPI behavior or calibration success. It does not prove why Sonnet 4.6 became unavailable. A separate Luna nonzero exit remains unresolved and should not be considered fixed by this change.
Acceptance Criteria