The same PDF, the same prompts, both models at temperature 0, five runs: Topic nodes 20, 37, 42, 2 and 17, "transfer wiedzy i mobilność" once a CriterionCategory with three children and once four Topic nodes with the sub-items baked into their titles, R2's competencies attached by HAS_CRITERION in one run and by REQUIRES/RECOMMENDS in the next, R4 a CompetencyCategory with 16 children in one run and a Topic with 4 in another. Every deterministic pass added narrows the damage a bad run does, none makes two runs agree.
A question that works on today's graph can stop working after the nightly refresh re-extracts a page, and no test can pin an answer that depends on which label or relationship type the model chose.
I propose: Not more prompt text. Constrain the model's choices: close the relationship vocabulary the way label_vocabulary closed the labels, and map drift to a canonical type, for a known document shape (a heading with enumerated rows) build the category -> item edges deterministically from the rows the completeness check already extracts, and let the model supply only titles and contexts. Measure with a fixed-page snapshot test that diffs label and relationship-type counts between two runs.
The same PDF, the same prompts, both models at temperature 0, five runs:
Topicnodes 20, 37, 42, 2 and 17, "transfer wiedzy i mobilność" once aCriterionCategorywith three children and once fourTopicnodes with the sub-items baked into their titles, R2's competencies attached byHAS_CRITERIONin one run and byREQUIRES/RECOMMENDSin the next, R4 aCompetencyCategorywith 16 children in one run and aTopicwith 4 in another. Every deterministic pass added narrows the damage a bad run does, none makes two runs agree.A question that works on today's graph can stop working after the nightly refresh re-extracts a page, and no test can pin an answer that depends on which label or relationship type the model chose.
I propose: Not more prompt text. Constrain the model's choices: close the relationship vocabulary the way
label_vocabularyclosed the labels, and map drift to a canonical type, for a known document shape (a heading with enumerated rows) build the category -> item edges deterministically from the rows the completeness check already extracts, and let the model supply only titles and contexts. Measure with a fixed-page snapshot test that diffs label and relationship-type counts between two runs.