Skip to content

Latest commit

 

History

36 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

gfw-ops

Python versions Last release

Reusable operational data engineering tools for GFW projects.

Features:

  • ✅ sharded-to-partitioned — migrates BigQuery sharded tables to partitioned tables.
  • ✅ bq-to-parquet — exports BigQuery tables to Parquet files on GCS.

Introduction

GFW maintains two shared Python repositories:

  • gfw-common — a pure library of reusable components (Beam transforms, BigQuery helpers, CLI framework, etc.) imported by pipeline repos.
  • pipe-* repos — individual pipeline applications tied to a specific data domain.

gfw-ops fills the gap between these two: it is the home for operational data engineering tools that are generic across all projects but are applications, not library code. Tools here run as config-driven jobs — triggered from Airflow, Cloud Build, or the command line — without requiring changes to any pipeline-specific repository.

Usage

Using the CLI

Write instructions on how to use the CLI of the application here.

Config file example

Optional. Provide an example of an input configuration file.

How to Contribute

Please read the guidelines in CONTRIBUTING.md.

Implementation details

Optional. This section is for describing implementation details, primarily for developers.

TBC.

Most relevant modules

Optional. Use this section to describe the most important modules of your application.

Example:

Module Description
cli.py Defines the application CLI.

About

Reusable operational data engineering tools for GFW projects.

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages