Reusable operational data engineering tools for GFW projects.
Features:
- ✅
sharded-to-partitioned— migrates BigQuery sharded tables to partitioned tables. - ✅
bq-to-parquet— exports BigQuery tables to Parquet files on GCS.
GFW maintains two shared Python repositories:
gfw-common— a pure library of reusable components (Beam transforms, BigQuery helpers, CLI framework, etc.) imported by pipeline repos.pipe-*repos — individual pipeline applications tied to a specific data domain.
gfw-ops fills the gap between these two:
it is the home for operational data engineering tools that are generic across all projects
but are applications, not library code.
Tools here run as config-driven jobs — triggered from Airflow,
Cloud Build, or the command line
— without requiring changes to any pipeline-specific repository.
Write instructions on how to use the CLI of the application here.
Optional. Provide an example of an input configuration file.
Please read the guidelines in CONTRIBUTING.md.
Optional. This section is for describing implementation details, primarily for developers.
TBC.
Optional. Use this section to describe the most important modules of your application.
Example:
| Module | Description |
|---|---|
| cli.py | Defines the application CLI. |