Skip to content

Extraction is not reproducible run to run #104

Description

@Dominik-Galus

The same PDF, the same prompts, both models at temperature 0, five runs: Topic nodes 20, 37, 42, 2 and 17, "transfer wiedzy i mobilność" once a CriterionCategory with three children and once four Topic nodes with the sub-items baked into their titles, R2's competencies attached by HAS_CRITERION in one run and by REQUIRES/RECOMMENDS in the next, R4 a CompetencyCategory with 16 children in one run and a Topic with 4 in another. Every deterministic pass added narrows the damage a bad run does, none makes two runs agree.

A question that works on today's graph can stop working after the nightly refresh re-extracts a page, and no test can pin an answer that depends on which label or relationship type the model chose.

I propose: Not more prompt text. Constrain the model's choices: close the relationship vocabulary the way label_vocabulary closed the labels, and map drift to a canonical type, for a known document shape (a heading with enumerated rows) build the category -> item edges deterministically from the rows the completeness check already extracts, and let the model supply only titles and contexts. Measure with a fixed-page snapshot test that diffs label and relationship-type counts between two runs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions