Add convert.parse object parsing wrangle - #1187
Conversation
- Introduce a new `convert.parse` recipe wrangle that parses JSON, Python literals, and YAML-like object text into JSON-compatible Python values. The implementation adds expected-type validation, per-column defaults, missing-value handling, normalization for numpy-backed objects, and safer YAML scalar resolution. - Eventually could replace from_yaml, from_json - Tests cover valid inputs, quoted structures, defaults, type mismatches, invalid values, and multi-column behavior.
There was a problem hiding this comment.
🟡 Changes recommended
Default normalization, schema validation, and YAML alias resource-safety issues remain unresolved.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Adds convert.parse for converting JSON, Python literals, and YAML-like text into JSON-compatible Python values.
Changes:
- Adds parsing, normalization, expected-type validation, and fallback handling.
- Adds comprehensive recipe-level tests for supported inputs and edge cases.
File summaries
| File | Description |
|---|---|
wrangles/recipe_wrangles/convert.py |
Implements convert.parse and its schema. |
tests/recipes/wrangles/test_convert.py |
Tests parsing, defaults, validation, and multiple columns. |
Review details
Suppressed comments (1)
wrangles/recipe_wrangles/convert.py:645
- The invalid/type-mismatch fallback also skips
_normalize_json_compatible, so NumPy-backed or otherwise unsupported defaults can escape unchanged. Apply the same normalization as successful parsed values so every return path honors the JSON-compatible result contract.
return _copy.deepcopy(col_default)
- Files reviewed: 2/2 changed files
- Comments generated: 3
- Review effort level: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| _convert_value(value) | ||
| for value in df[input_column] | ||
| ] | ||
|
|
There was a problem hiding this comment.
@ebhills
Assigning the parsed results as a plain list lets pandas infer a float dtype when the column contains both integers and floats. This silently corrupts large integers: parsing ["9007199254740993", "1.5"] produces [9007199254740992.0, 1.5].
Reproduced through a recipe and confirmed by rendering value={{ parsed }} as text before exporting to Excel, so this is not an Excel display issue.
Please assign the results using an explicitly object-typed Series with index=df.index, and add regression tests for a large integer alongside a float or a missing value.
There was a problem hiding this comment.
Fixed in f30e6be1. convert.parse now assigns an explicitly object-typed Series with index=df.index, preserving the exact Python integer 9007199254740993 alongside either a float or a missing value.
Regression coverage: test_large_integer_precision_with_mixed_scalars checks exact values, Python types, and a non-default row index. test_large_integer_precision_survives_recipe_rendering verifies the recipe path and exact value={{ parsed }} text for both mixed-column cases. These four cases failed before the fix and now pass; all 181 focused conversion/schema/DataFrame tests passed locally. Fresh CI is running.
Recommended disposition: Comment only
Next steps
- PR assignee: After CI passes, re-request review from
mborodii-progon Add convert.parse object parsing wrangle #1187. - Reviewer: Verify the precision and index regressions, resolve this conversation, and submit a fresh approval once both fixes are verified.
| f"Result is not a valid {col_expected}" | ||
| ) | ||
| return result | ||
| except (TypeError, ValueError) as error: |
There was a problem hiding this comment.
@ebhills Deeply nested input raises RecursionError, which this handler does not catch. For example, parsing "[" * 1100 + "0" + "]" * 1100 with default={} aborts the recipe instead of returning the configured fallback.
Please handle recursion-limit failures as conversion failures, or enforce a depth limit that raises a handled exception. Add regression coverage confirming that excessive nesting returns the configured default, while the same input without a default raises a contextual ValueError.
read:
- test:
rows: 1
values:
original: placeholder
wrangles:
- create.jinja:
output: original
template:
string: '{{ "[" * 1100 }}0{{ "]" * 1100 }}' - convert.parse:
input: original
output: parsed
default: {}
There was a problem hiding this comment.
Fixed in f30e6be1. The conversion handler now catches RecursionError: excessively nested input returns an independent copy of the configured default, or raises the contextual ValueError when no default is supplied. Error formatting also handles deeply nested materialized objects safely, and an excessively nested explicit default raises a contextual invalid-default error.
Regression coverage: test_excessive_nesting_uses_default and test_excessive_nesting_without_default_has_context cover JSON, YAML, and materialized objects through both the DataFrame API and recipe runner. test_excessively_nested_default_is_rejected covers default validation. These 13 cases failed before the fix and now pass; all 181 focused conversion/schema/DataFrame tests passed locally. Fresh CI is running.
Recommended disposition: Comment only
Next steps
- PR assignee: After CI passes, re-request review from
mborodii-progon Add convert.parse object parsing wrangle #1187. - Reviewer: Verify fallback and contextual-error behavior for excessive nesting, resolve this conversation, and submit a fresh approval once both fixes are verified.
Adds
convert.parsefor cells containing JSON, Python literals, YAML-like structures, or accidentally quoted objects. It produces JSON-compatible values with explicit expected-type and fallback handling while preserving the behavior of the four existing JSON/YAML conversion wrangles.Linked issue
Closes #1189
What changes
expected: any,dictionary,list, orscalar, including per-column categories and defaults. Missing cells remain empty when no default is supplied.ValueError; mutable defaults are copied independently for each row.ValueError. Deeply nested defaults are rejected with a contextualValueError, and error reporting remains safe for deeply nested materialized objects.00123,12:34, and0xFFremain strings, while ordinary decimal and scientific-notation numbers remain numeric.expectedseparately in the generated recipe schema.Scope is limited to
wrangles/recipe_wrangles/convert.pyandtests/recipes/wrangles/test_convert.py.How it was verified
Validated commit
f30e6be1, which includes currentmain(0df6c569).tests/recipes/wrangles/test_convert.py,tests/recipes/wrangles/test_main.py::TestWrangleSchema, andtests/test_dataframe.py. Credentials were removed from the test process and network access was blocked.expectedvalues.expectedvalues were rejected.git diff --checkpassed. Implementation comparison confirmed thatfrom_json,from_yaml,to_json, andto_yamlremain unchanged.f30e6be1. Local checks do not establish completed CI or deployment.Compatibility and risk
This is an additive wrangle. Existing converters and their callers retain their current behavior. No credentials or external services are required by
convert.parse, and no deployment is included.The new parser deliberately rejects YAML anchors, aliases, explicit tags, and non-JSON-compatible values/defaults. YAML-only numeric forms are retained as text. Its output columns retain object dtype to preserve Python value types and integer precision; excessive nesting follows the same default/error contract as other conversion failures. These rules are documented in its schema docstring and covered by regression tests.
Future direction:
convert.parsecould eventually replace some uses offrom_yamlorfrom_json; migration or replacement of those wrangles remains outside this PR, as does recursively parsing serialized strings nested inside objects.Rollback: remove any newly introduced
convert.parserecipe usage, then revert this PR. Existing converters require no migration.The PR is Ready with changes requested. The two current review findings are fixed in
f30e6be1and await reviewer verification; the earlier three conversations are resolved. After fresh CI passes and both current threads have fixing-commit replies, the human assignee should re-request review frommborodii-prog. The reviewer should verify the fixes, resolve the two conversations, and submit a fresh approval.Ready-for-review checklist
mainand has no merge conflictsNo release milestone is selected.
See the pull request workflow.