Problem
When a test grader is configured with Negate, AITestRunner flips whatever the grader returns (Umbraco.AI/src/Umbraco.AI.Core/Tests/AITestRunner.cs, the if (grader.Negate) block). Graders report their own errors as ordinary failed results, so a grader that couldn't grade is flipped into a pass.
Examples:
llm-judge returns Passed = false with "LLM judge evaluation failed: ..." when the chat call fails or the judgment can't be parsed. With Negate on, the test passes.
- The new experimental
decision-judge grader (Decision capability, not merged yet) fails with "Decision is turned off ..." or "Decision judge evaluation failed: ..." when it can't run. With Negate on, those become passes too.
Grader exceptions caught by the runner itself are not negated (the catch records a failure), so only errors a grader handles internally are affected.
Expected behavior
A grader that could not produce a verdict (provider error, misconfiguration, feature turned off, unparseable response) should fail the test whether or not Negate is set. Negate should only invert a real verdict.
Notes
- The runner can't tell an error result from a real "no" today:
AITestGraderResult has no error marker, so this likely needs one (or graders throwing on errors and the runner handling it).
- Applies to both active lines (v18 and v17).
- Found while building the Decision test grader (
v18/feature/decision-evaluators); its docs currently warn about this limitation.
Problem
When a test grader is configured with Negate,
AITestRunnerflips whatever the grader returns (Umbraco.AI/src/Umbraco.AI.Core/Tests/AITestRunner.cs, theif (grader.Negate)block). Graders report their own errors as ordinary failed results, so a grader that couldn't grade is flipped into a pass.Examples:
llm-judgereturnsPassed = falsewith "LLM judge evaluation failed: ..." when the chat call fails or the judgment can't be parsed. With Negate on, the test passes.decision-judgegrader (Decision capability, not merged yet) fails with "Decision is turned off ..." or "Decision judge evaluation failed: ..." when it can't run. With Negate on, those become passes too.Grader exceptions caught by the runner itself are not negated (the
catchrecords a failure), so only errors a grader handles internally are affected.Expected behavior
A grader that could not produce a verdict (provider error, misconfiguration, feature turned off, unparseable response) should fail the test whether or not Negate is set. Negate should only invert a real verdict.
Notes
AITestGraderResulthas no error marker, so this likely needs one (or graders throwing on errors and the runner handling it).v18/feature/decision-evaluators); its docs currently warn about this limitation.