Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -424,3 +424,26 @@ Adding a feature to the tier matrix is an entry in `src/core/registry.ts`, not a
Licensed under [Apache 2.0](LICENSE). See [`NOTICE`](NOTICE) for attribution.

The probe, scoring, conformance tests, capability evals, benchmark, and CLI are original to llmprobe. The OpenAPI schemas under `schema/` (and the Zod validators generated from them) are derived from [openresponses](https://github.com/openresponses/openresponses) and the official [openai/openai-openapi](https://github.com/openai/openai-openapi) spec, retained under Apache 2.0 — full attribution in [`NOTICE`](NOTICE).

### Multilingual response collection

`--language-only` collects 100 original prompts: eight tasks in each of English,
Chinese, Thai, Japanese, Spanish, French, German, Korean, Arabic, Portuguese,
Russian and Italian, plus four cross-language tasks. Categories cover register,
idioms, pragmatics, grammar, localization, culture, creative writing and ambiguity.

```bash
npx llmprobe localhost:12345 --language-only --model MODEL \
--concurrency 4 --eval-max-tokens 500 --timeout 600 --save runs/language-MODEL.jsonl
```

This mode defaults to greedy sampling, thinking off, speculation off, concurrency
4 and a 500-token output cap. The OpenAI-compatible endpoint must support these
request controls. It writes a manifest, every complete response (including usage,
finish reason and request settings), errors, and a summary to a new JSONL file.
Existing logs are never overwritten. It does not upload data or assign automatic
language scores; evaluate saved answers offline. Scores should cover accuracy,
naturalness, register/style, nuance/cultural fit and task fulfillment, with equal
weight per language. Report cross-language cases separately, and disclose the
judge and any output truncation. These logs are distinct from normal probe report
JSON and cannot be passed to `--compare`.
40 changes: 40 additions & 0 deletions bin/llmprobe.ts
Original file line number Diff line number Diff line change
Expand Up @@ -94,7 +94,10 @@ import {
} from "../src/reasoning/index";
import { ALL_EVALS } from "../src/evals/index";

import { runLanguage } from "../src/language/index";

interface Args {
languageOnly?: boolean;
target?: string;
apiKey?: string;
model?: string;
Expand Down Expand Up @@ -241,6 +244,9 @@ function parseArgs(argv: string[]): Args {
args.bench = true;
args.benchOnly = true;
break;
case "--language-only":
args.languageOnly = true;
break;
case "--eval":
args.eval = true;
break;
Expand Down Expand Up @@ -453,6 +459,10 @@ Options:
On by default; --no-bench skips it
--bench-only Run only the benchmark — no conformance, evals, agentic
or fidelity. Surface discovery still runs; it is free.
--language-only Collect 100 multilingual responses for offline review (12 languages).
Requires --model and --save (new JSONL file). Defaults: concurrency 4,
max output 500, greedy sampling, thinking off. Uses --concurrency,
--eval-max-tokens and --timeout. Separate from scored probe reports.
--eval Reasoning accuracy: 92 questions from GPQA Diamond,
SuperGPQA, AIME 2025 and COMPSEC (informational, never
scored). Expensive on a thinking model: up to
Expand Down Expand Up @@ -1717,6 +1727,36 @@ async function main(): Promise<void> {

const root = normalizeRoot(args.target);
const apiKey = args.apiKey ?? process.env.LLMPROBE_API_KEY ?? "";
if (args.languageOnly) {
if (!args.model || !args.save)
throw new Error("--language-only requires --model and --save");
if (
args.sampling ||
args.upload ||
args.html ||
args.eval ||
args.benchOnly ||
args.budget ||
args.evalCases ||
args.evalQuestions
)
throw new Error(
"--language-only supports --model, --save, --concurrency, --eval-max-tokens, --timeout and --api-key; other run modes are not supported",
);
const result = await runLanguage({
root,
model: args.model,
save: args.save,
apiKey,
concurrency: args.concurrency,
maxTokens: args.evalMaxTokens,
timeoutMs: args.timeoutSec * 1000,
log,
});
console.log(JSON.stringify(result));
if (result.errors) process.exitCode = 1;
return;
}
const startedAt = Date.now();

log(
Expand Down
Loading