You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add explicit lookup.key and lookup.semantic operations to WranglesPY, with separate recipe contract definitions.
The existing public entry point is generic lookup; the saved model's variant determines whether execution uses key lookup or semantic search. Explicit operation names will make that distinction visible in Python calls, recipes, generated schemas, and downstream documentation.
This follows the discussion in Wrangles-Docs #34, which proposed implementing the lookup variants as callable operations, similar to the existing extract.ai and extract.custom distinction.
Desired Behavior
Expose these operations through Python, YAML recipes, and the Wrangles DataFrame accessor:
Operation
Saved model purpose
Stored model variant
lookup.key
lookup
key
lookup.semantic
lookup
embedding
Reuse the existing lookup execution path and saved models.
Keep generic lookup available for existing Python callers and recipes.
Validate that an explicitly selected operation is compatible with the supplied model. Make handling of missing and unknown variants explicit.
Use the existing execution model_id. A catalog ID must not be substituted for the selected saved-model ID.
Preserve supported output selection and renaming, complete-record output, row order, and existing lookup modes.
Define each operation's recipe contract through the existing schema-generation mechanism: required parameters, defaults, accepted input/output forms, mode constraints, and examples.
Document how n affects return shapes and which recipe modes support it.
Semantic models continue to use stored variant=embedding. The new operation name is lookup.semantic.
Example current / desired
In these examples, KEY_MODEL_ID and SEMANTIC_MODEL_ID are existing saved-model IDs. Define them as recipe variables or replace the placeholders with actual IDs.
The example models expose their value under Column1, matching the manual Excel fixtures. Replace this with the actual stored field name when using another model, such as Output. The left side of an output mapping selects the model field; the right side names the result column.
For a model containing ABC-001 → Bolt and ABC-002 → Nut, the selected field should contain Bolt and Nut, with duplicate input rows and their order preserved.
For the test fixture, descriptions such as steel bolt and steel bolt eight millimetre should match the stored bolt description. The selected value and score should appear in separate columns.
Complete-record and multiple-match output should also remain available. For example:
With three available matches, Match 1, Match 2, and Match 3 should contain match dictionaries. The lookup backend remains responsible for ranking and scores.
Both explicit names resolve and execute through Python, recipes, and the Wrangles DataFrame accessor.
Each operation validates model purpose and the agreed variant rules, with useful errors for incompatible models.
Existing generic lookup calls and saved recipes retain their behavior.
Generated recipe schemas contain separate definitions for lookup, lookup.key, and lookup.semantic.
Contract definitions match actual parameters, defaults, output forms, and supported mode combinations.
Tests cover scalar/list inputs, selected fields and output renaming, complete records, duplicate rows, empty inputs, supported lookup modes, multiple matches, and invalid model variants.
Regression tests compare legacy and explicit operations using the same model metadata and backend responses.
Manual Excel checks cover key and semantic execution with the same input and model when comparing against legacy lookup; score differences and untested cases are recorded explicitly.
Examples use actual saved-model field names and execution IDs, with known limitations documented.
Draft behavior to confirm during review
The current draft requires exact stored variants: key for lookup.key and embedding for lookup.semantic. Missing/null, unknown, and literal semantic variants are rejected by the explicit names. Generic lookup retains its existing handling.
The new recipe operations reject a non-null n with by_dataframe or by_matrix; multiple-match requests use by_row.
These are proposed contract choices for review, not changes to existing generic lookup behavior.
Scope
This issue covers the WranglesPY operations, recipe contract definitions, and their tests. Separate documentation pages and Registry integration should be coordinated with the Wrangles-Docs work.
It does not require new database tables, catalog allocation changes, model migrations, a new lookup API endpoint, or a new semantic scoring algorithm.
Description
Add explicit
lookup.keyandlookup.semanticoperations to WranglesPY, with separate recipe contract definitions.The existing public entry point is generic
lookup; the saved model's variant determines whether execution uses key lookup or semantic search. Explicit operation names will make that distinction visible in Python calls, recipes, generated schemas, and downstream documentation.This follows the discussion in Wrangles-Docs #34, which proposed implementing the lookup variants as callable operations, similar to the existing
extract.aiandextract.customdistinction.Desired Behavior
Expose these operations through Python, YAML recipes, and the Wrangles DataFrame accessor:
lookup.keylookupkeylookup.semanticlookupembeddinglookupavailable for existing Python callers and recipes.model_id. A catalog ID must not be substituted for the selected saved-model ID.naffects return shapes and which recipe modes support it.Semantic models continue to use stored
variant=embedding. The new operation name islookup.semantic.Example current / desired
In these examples,
KEY_MODEL_IDandSEMANTIC_MODEL_IDare existing saved-model IDs. Define them as recipe variables or replace the placeholders with actual IDs.The example models expose their value under
Column1, matching the manual Excel fixtures. Replace this with the actual stored field name when using another model, such asOutput. The left side of an output mapping selects the model field; the right side names the result column.Current generic key lookup:
Desired explicit key lookup:
For a model containing
ABC-001 → BoltandABC-002 → Nut, the selected field should containBoltandNut, with duplicate input rows and their order preserved.Desired explicit semantic lookup:
For the test fixture, descriptions such as
steel boltandsteel bolt eight millimetreshould match the stored bolt description. The selected value and score should appear in separate columns.Complete-record and multiple-match output should also remain available. For example:
With three available matches,
Match 1,Match 2, andMatch 3should contain match dictionaries. The lookup backend remains responsible for ranking and scores.Python usage:
Acceptance criteria
lookupcalls and saved recipes retain their behavior.lookup,lookup.key, andlookup.semantic.Draft behavior to confirm during review
keyforlookup.keyandembeddingforlookup.semantic. Missing/null, unknown, and literalsemanticvariants are rejected by the explicit names. Genericlookupretains its existing handling.nwithby_dataframeorby_matrix; multiple-match requests useby_row.These are proposed contract choices for review, not changes to existing generic lookup behavior.
Scope
This issue covers the WranglesPY operations, recipe contract definitions, and their tests. Separate documentation pages and Registry integration should be coordinated with the Wrangles-Docs work.
It does not require new database tables, catalog allocation changes, model migrations, a new lookup API endpoint, or a new semantic scoring algorithm.