Skip to content

[AURON #2558] Support Iceberg _partition metadata in native scans - #2566

Open
Sigma-Ma wants to merge 2 commits into
apache:masterfrom
Sigma-Ma:Auron-2558-support-iceberg-partition-metadata-column-in-native-scans
Open

Sigma-Ma wants to merge 2 commits into
apache:masterfrom
Sigma-Ma:Auron-2558-support-iceberg-partition-metadata-column-in-native-scans

Conversation

@Sigma-Ma

@Sigma-Ma Sigma-Ma commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Closes #2558

Rationale for this change

Queries projecting _partition currently fall back to Spark even when the Iceberg COW scan is otherwise eligible for native execution.

What changes are included in this PR?

  • Materialize _partition through the existing per-task constant path using Iceberg's unified partition type and PartitionUtil.constantsMap.
  • Preserve Iceberg partition field IDs when selecting projected fields, since Spark may renumber nested metadata fields.
  • Serialize struct constants from their existing Catalyst representation.
  • Add Parquet and ORC integration coverage for empty and unpartitioned tables, NULL values, identity/bucket/day transforms, partition evolution, field projections, and fallback.

The scope covers regular COW scans. _pos, delete-file handling, and _partition in changelog scans remain outside this change.

Are there any user-facing changes?

Eligible COW queries projecting _partition or its fields can use NativeIcebergTableScan. No new configuration is required.

How was this patch tested?

  • ./dev/reformat --check.
  • ./build/mvn -B -ntp -Ppre,spark-3.5,scala-2.12,iceberg-1.10.1 -pl thirdparty/auron-iceberg -am -DskipTests clean install.
  • ./build/mvn -B -ntp -Ppre,spark-3.5,scala-2.12,iceberg-1.10.1 -pl thirdparty/auron-iceberg -DskipBuildNative -Dsuites=org.apache.auron.iceberg.AuronIcebergIntegrationSuite test.

Was this patch authored or co-authored using generative AI tooling?

  • Yes
  • No

Generated-by: OpenAI Codex (GPT-6)

ASF guidance: https://www.apache.org/legal/generative-tooling.html

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The implementation follows Iceberg’s constant-materialization semantics and includes broad integration coverage.

0 open findings

What changed in this PR

Adds native Iceberg COW scan support for _partition metadata.

Changes:

  • Materializes unified partition metadata via Iceberg task constants.
  • Preserves partition field IDs and serializes Catalyst struct literals.
  • Adds Parquet/ORC integration and fallback coverage.
File Description
AuronIcebergIntegrationSuite.scala Tests partition metadata and fallback scenarios.
NativeIcebergTableScanExec.scala Serializes existing Catalyst partition values.
IcebergScanSupport.scala Materializes _partition task constants.
AuronIcebergSourceUtil.scala Builds projected unified partition types with Iceberg IDs.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support the Iceberg _partition metadata column in native scans

2 participants