Repository navigation
feat(iceberg): replay-safe immutable history ingestion - #1077
Conversation
Validate and align incoming data before staging writes. Use one Iceberg transaction for every delete batch and the replacement append, and record post-commit evidence only after publication. Add local catalog acceptance tests for multiple delete batches, rollback, replay, real commit conflicts, resource evidence, and DLT ingestion with ref isolation. Closes #1064
Expose strict, additive, and drop_extra policies on ingestion and storage writes. Validate Arrow batches before mutation and stage nullable additions with data in the existing transaction. Reconcile compatible concurrent additions, preserve unrelated provider defaults, and document the intentional compatibility change.
Compare explicit entity/version identity and payload policy across a complete staged batch. Bind policy and schema additions with snapshot-guarded insertion, retry complete operations on definite conflicts, and reconcile ambiguous outcomes without claiming counts. Expose optional provider support, ingestion evidence, real-catalog and service acceptance tests, and history configuration and migration documentation.
|
Note This drawing shows
Nothing flagged · reviewed Architecture Inside the changed components — 2 viewsComponent view — dlt Ingestion Pipeline Internal decorator, executor, and helper components within phlo-dlt. Component view — Iceberg History Engine Internal resource adapter and history writer within phlo-iceberg. Data flow The other flows — 1 sequence
View
Tip Open a diagram on the canvas, then press W or click play to walk through the change one step at a time 🪧 More tips
Thanks for using PR Lens! It's built by Coldtea, free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. |
|
Reviewed 1. Lookup is filtered on the version key only, so read cost scales with the table rather than the batch
Under the identity scheme in #1066, 2.
Checked and deliberately not flagged. The central concurrency claim holds. In PyIceberg 0.11.0/0.11.1 an ordinary append staged from a branch always emits Validation not reproduced here. I ran no checks in this review; the |
Push exact composite identity predicates into snapshot-pinned scans in groups of at most 1000 pairs. Report distinct matching, missing, and conflicting version counts after ambiguous commits rather than counting duplicate observations as stored rows. Exercise real scan results, crossed identities, batching, landed and unlanded outcomes, reconciliation conflicts, and public ingestion evidence. Document distinct-version reconciliation semantics.
Forward explicit migration policies across every chunk and Kafka policies only to opted-in stores. Formalise the optional provider contract and validate advertised signatures before writes while preserving legacy defaults. Project unified carrier business columns before strict writes, exercise real materialisation and local-catalog wrapper behavior, and document configuration and compatibility boundaries.
Preserve the optional schema-policy signature negotiation and validate the selected history writer before table creation. Retain both published histories and add history-specific contract coverage.
Preserve the approved history-only delta on top of the schema-policy squash and main export/service changes. Verify the resolved tree exactly equals main plus the previously approved history patch.
Scope
Closes #1066. Stacked on #1075 (
feat/1065-iceberg-schema-policy, exact base e47b0b5), which follows #1064. This diff contains only immutable-history support, evidence, tests, and documentation; no merge or schema-policy reimplementation.Behaviour
merge_strategy="history"through supported DLT ingestion and an optional provider capability; unsupported providers reject before table creation or writes.The uniqueness guarantee covers cooperating history writers on the same table/ref. Ordinary writes and property-only migration must not race history ingestion. Batches must fit memory. Domain identity generation, latest-version ordering, and live migration are outside scope.
Validation
make setup: passed.make check: passed; 5,932 passed, 4 skipped, 229 deselected. An initial shallow-history header-test failure disappeared after fetching full repository history; the complete baseline was rerun successfully.uv run --locked pytest packages/phlo-iceberg/tests packages/phlo-dlt/tests -m "not integration" --tb=short -q: 424 passed, 7 deselected, including 58 new history cases.uv run --locked python scripts/run_integration.py: 111 passed, no skips; includes real Nessie/MinIO history replay, conflict rejection, and injected concurrent commit retry. An initial test-only catalog-instance injection mistake was corrected and the full lane rerun.make test-core-regression: 360 passed, 2,381 deselected.make docs-build: passed, 849 static pages; inspected rendered guide and API reference, including scrolling the code example to its final lines.git diff --check: passed.Real-catalog cases cover committed and in-batch conflicts, composite identities, late versions, null/NaN payloads, supplied hashes, multifile input, bounded retries, concurrent empty/nonempty tables, additive races, competing policies, unknown landed/unlanded outcomes, and evidence. The public decorated ingestion test stages actual DLT data.
No dependency changes. PyIceberg 0.11.0 (declared minimum) and locked 0.11.1 both emit the snapshot-head assertion, including an absent initial head.