Skip to content

Proposal: replace Sonnet 4.6 with Sonnet 5 in baseline-equivalence CI #2899

Description

@jkim323

Issue Description

Propose replacing claude-sonnet-4.6 with claude-sonnet-5 in the baseline-equivalence model list while retaining gpt-5.6-luna and the existing validation gates.

During investigation of PR #2894, Sonnet 4.6 was absent from the model catalog returned under the CI configuration, and session creation failed with model-unavailable. A subsequent isolated test targeting Sonnet 5 passed across all four tested Vally/Copilot SDK combinations, including the repository’s current versions.

Proposed Change

Update Resolve-ModelList in scripts/evals/Invoke-BaselineEquivalence.ps1 so the calibration and ci tiers use:

  • gpt-5.6-luna
  • claude-sonnet-5

Update associated model-selection tests, help text, and current documentation. Preserve historical results and references that describe earlier runs.

Keep the comparison judge, retry policy, execution-health checks, evidence-completeness checks, invariants, and tier-specific gating behavior unchanged. No Vally or Copilot SDK dependency change is proposed.

This retains cross-model coverage rather than reducing the suite to Luna-only.

Additional Context

The isolated harness changed only the requested model ID. It used the existing CI credential configuration, a fixed arithmetic prompt, no permitted tools, and the same CLI version.

Vally Copilot SDK Sonnet 4.6 Sonnet 5
0.14.0 1.0.11 Model unavailable Pass
0.15.0 1.0.9 (current) Model unavailable Pass
0.14.0 1.0.9 Model unavailable Pass
0.15.0 1.0.11 Model unavailable Pass

All Sonnet 5 cells reported catalog presence, the expected response, zero tool-permission requests, and successful cleanup. Recorded per-cell dependency-lock and CLI fingerprints matched the earlier Sonnet 4.6 test.

Evidence:

The smoke establishes basic availability, not full RPI behavior or calibration success. It does not prove why Sonnet 4.6 became unavailable. A separate Luna nonzero exit remains unresolved and should not be considered fixed by this change.

Acceptance Criteria

  • Maintainers approve the proposed model replacement.
  • Calibration and CI select Luna and Sonnet 5; devloop override behavior remains unchanged.
  • Relevant tests, help text, and current documentation reflect the new selection.
  • Existing validation gates, judge configuration, and retry behavior remain unchanged.
  • Targeted model-selection and runner tests pass.
  • Full calibration is run under the CI configuration, with complete Sonnet 5 execution evidence and results assessed under existing tier rules.
  • Remaining Luna or other failures are tracked separately; smoke success is not reported as full CI acceptance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

evalsmaintenanceMaintenance work, no version bumppriority-2High priority, address soontestingTest infrastructure and test files

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions