Conversation
Ruiying-Ma
added a commit
that referenced
this pull request
Sep 26, 2026
) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RhSArH6Zi4Gfka4XRrHJpr
Collaborator
|
Hi @Legend398! Thanks for the submission and the detailed traces. We have added your results to the leaderboard. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
DVIndex — Astra/high with Jev classification — 87.68% local Pass@1
Please review this 270-record DVIndex submission and its disclosed complete-question reruns. Local scoring gives 238/270 raw, 0.8768498168 dataset-macro Pass@1, pending your validation.
Reruns and interruptions
All five trials were rerun and replaced for AGNews Q2/Q3/Q4 and CRM Q2/Q7/Q8/Q12. All 35 rerun outcomes are included; no individual result was selected by correctness. Classification reruns repair a prepared-input launch integration error. CRM reruns use revised dataset method guidance and are disclosed as a changed prompt configuration. The other 235 original pairs remain unchanged.
The original run was paused/resumed across 19–23 September. Real timestamps and failed/cancelled outcomes are retained. One previously performed isolated Yelp Q7 run 2 retry is supplied as supplemental evidence but is not selected; the main result counts its original cancellation as a non-pass. Original superseded recordings are also included for audit.
Files
leaderboard_submissions/dvindex.jsonis the answer file. The complete evidence archive (download page) contains full traces, original evidence, per-row Jev receipts, prompts, timing, replacement-map.json, configuration fingerprints and local verification. Nothing in this draft asserts that the replacement policy or prompts have already been approved. Please advise if additional evidence or a different treatment is required.Local results
Scored against unchanged official validators at aae9730; current-main material equality was checked on 23 September. Every selected answer matches its recorded final answer. No manual score overrides or answer rewrites were applied.