Skip to content

feat: collect multilingual responses for offline comparison - #5

Open
beamivalice wants to merge 1 commit into
ddalcu:mainfrom
beamivalice:codex/multilingual-comparison
Open

beamivalice wants to merge 1 commit into
ddalcu:mainfrom
beamivalice:codex/multilingual-comparison

Conversation

@beamivalice

Copy link
Copy Markdown

Add --language-only to collect comparable multilingual model responses for offline review. The suite includes 100 original prompts: eight matched task categories in each of 12 languages, plus four cross-language cases.

The mode requires --model and a new --save JSONL path. It records a manifest with the prompt hash and review rubric, complete responses and request settings, errors, and a summary. Defaults are concurrency 4, a 500-token output cap, greedy sampling, and thinking/speculation disabled. Existing logs are never overwritten, and request failures produce a nonzero exit status.

The rubric defines five scoring dimensions, equal weighting per language, paired comparison guidance, and separate reporting for cross-language cases, missing answers, and truncation. This is response collection with an offline review protocol; it does not automatically score language quality, upload results, or feed the existing --compare report workflow. Endpoints must support the documented request controls.

Validation against current upstream main:

  • npm test: 468 tests passed across 36 files, including prompt coverage, bounded concurrency, complete Unicode responses, error recording, and overwrite protection.
  • npm run typecheck: passed.
  • npx prettier --check . --ignore-unknown: passed.
  • CLI build and --help smoke check: passed.

No live model evaluation was run for this PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant