Repository navigation
PIPELINE-4277: Add GCS Parquet compaction pipeline - #9
Merged
Merged
Conversation
tomaslink
force-pushed
the
feature/compact-parquet
branch
7 times, most recently
from
July 3, 2026 15:01
c98c67a to
d9bfeea
Compare
Compacts small hive-partitioned Parquet files on GCS into larger files using a staging-based swap to keep the original path intact. Adds pyarrow and gcsfs dependencies. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
tomaslink
force-pushed
the
feature/compact-parquet
branch
2 times, most recently
from
July 6, 2026 18:27
0baf2c6 to
6b5e32c
Compare
tomaslink
force-pushed
the
feature/compact-parquet
branch
4 times, most recently
from
July 7, 2026 23:20
bc7757c to
c6e3548
Compare
andres-arana
requested changes
Jul 9, 2026
andres-arana
left a comment
There was a problem hiding this comment.
I think the big one is the potential dangerous delete. Let's brainstorm a solution for that here. Everything else is minor (but easy to fix).
tomaslink
force-pushed
the
feature/compact-parquet
branch
from
July 10, 2026 13:49
c6e3548 to
09fa16a
Compare
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
tomaslink
force-pushed
the
feature/compact-parquet
branch
from
July 10, 2026 14:01
09fa16a to
b11e11e
Compare
A crash partway through deleting original blobs left the live partition
with a partial remnant that the old resume check ("staging present, source
empty") couldn't distinguish from a fresh partition, causing the correct
staged replacement to be wiped and only the survivors recompacted.
Record the exact list of originals to delete in a manifest before starting
the delete, so a resumed run can finish from that authoritative list
regardless of how far the previous attempt got. Raise instead of guessing
when staging is non-empty, source is empty, and no manifest exists, since
that state is unverifiable and should not occur post-migration.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
https://globalfishingwatch.atlassian.net/browse/PIPELINE-4277
Summary
Adds two new pipelines for working with hive-partitioned Parquet files on GCS:
compact-parquetandbenchmark-parquet.compact-parquet
Compacts small Parquet files within each date partition into larger files using DuckDB. Operates in two modes:
--gcs-staging-path): writes compacted output to a separate path, leaving source files untouched — useful when you want both uncompacted and compacted versions available via separate external tables.DuckDB memory and thread count are configurable via
--memory-limitand--threads.benchmark-parquet
Runs configurable queries against one or more BigQuery tables and prints a comparison of bytes processed, slot milliseconds, and wall-clock time. Supports
SELECT *and an hourly aggregation query, with query cache disabled so each run reflects real scan cost. Designed to compare native BigQuery tables against Parquet external tables.Tests
Both pipelines have 100% test coverage. Compact-parquet tests cover the full compaction flow, swap/copy modes, interrupted swap resume, leftover staging cleanup, and skip logic — without patching private methods.