A learning project for crawling news, storing data in Google BigQuery, and running crawlers on Cloud Functions (simple) and Cloud Run (headless browser).
Disclaimer: This project is for educational and learning purposes only. If any content or use infringes your rights, please contact the author for immediate removal.
- Serverless: Cloud Functions for on-demand runs
- Storage: BigQuery with date partitioning and SQL access
- Deduplication: In-memory link cache at startup + partition-aware queries
- Concurrency: Multiple sources run in parallel
- Scheduling: Cloud Scheduler (e.g. every 10–30 minutes)
- Low cost: Roughly $4–7/month for the simple setup
- Runtime: Python 3.11
- Compute: Google Cloud Functions (Gen 2), Cloud Run (browser)
- Database: Google BigQuery (partitioned by date)
- Scheduling: Google Cloud Scheduler
- Dependencies: BeautifulSoup4, requests, google-cloud-bigquery; Playwright + Firefox for browser crawlers
news-google/
├── main.py # Dual entry: crawl_news (simple) + crawl_news_browser (headless)
├── Makefile # Local run & deploy (make run / deploy / deploy-browser)
├── requirements.txt # Simple crawler deps
├── requirements-browser.txt # Headless browser deps (playwright, etc.)
├── config.yaml # Config (GCP, concurrency)
├── scrapers/
│ ├── base_scraper.py
│ ├── simple/ # Simple crawlers (requests/BeautifulSoup, no browser) → Cloud Functions
│ └── browser/ # Headless browser crawlers (Playwright) → Cloud Run only
│ └── __init__.py # SCRAPER_REGISTRY_BROWSER
├── utils/
└── deploy/
├── deploy.sh # Deploy crawl-news (simple) → Cloud Functions
├── deploy_cloudrun_browser.sh # Browser crawler → Cloud Run (+ Scheduler)
├── setup_scheduler.sh # Schedule crawl-news
└── create_bigquery_table.sql
└── Dockerfile.firefox # Cloud Run image (Firefox only)
Two entry points (separate dependencies)
| Type | Entry | Deps | Deploy target |
|---|---|---|---|
| Simple crawlers | crawl_news |
requirements.txt | Cloud Functions |
| Headless browser | crawl_news_browser |
requirements-browser + browser binary | Cloud Run (Dockerfile.firefox) |
Headless crawlers need Playwright’s browser binary → deploy via Cloud Run + Dockerfile.firefox.
- Install Google Cloud SDK and log in:
gcloud auth login gcloud auth application-default login
- Set env and config:
Edit
export GCP_PROJECT_ID="your-project-id" export GCP_REGION="us-central1"
config.yamland replaceyour-project-idwith your GCP project ID.
bq query --use_legacy_sql=false < deploy/create_bigquery_table.sqlUsing Makefile (recommended)
make deploy # Simple crawler → Cloud Functions
make deploy-browser # Browser crawler → Cloud Run + Scheduler
make deploy-all # BothOr run scripts directly
sh deploy/deploy.sh
sh deploy/deploy_cloudrun_browser.sh- Simple crawler schedule (e.g. every 10 min):
./deploy/setup_scheduler.sh - Browser crawler: Scheduler is set up by
deploy_cloudrun_browser.sh(e.g. every 30 min).
Trigger via HTTP
curl -X POST -H "Content-Type: application/json" \
-d '{"sources": "all"}' \
https://REGION-PROJECT.cloudfunctions.net/crawl-newsLocal run (from repo root; GCP auth required for non-test)
- Simple crawlers:
make runoruv run python main.py - Browser crawlers:
make install-browseronce, thenmake run-browseroruv run python main.py browser
Local default istest=True(no BigQuery writes).
Table is partitioned by pub_date; queries must include a pub_date filter.
SELECT title, link, source, pub_date, crawled_at
FROM `YOUR_PROJECT_ID.news_project.news_articles`
WHERE DATE(pub_date) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
ORDER BY crawled_at DESC
LIMIT 20;
SELECT source, COUNT(*) AS cnt
FROM `YOUR_PROJECT_ID.news_project.news_articles`
WHERE DATE(pub_date) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
GROUP BY source
ORDER BY cnt DESC;- Simple: Add a module under
scrapers/simple/, extendBaseScraper, register inmain.py→SCRAPER_REGISTRY, thenmake deploy. - Browser: Add a module under
scrapers/browser/, extendBaseBrowserScraperand implement_run_impl(), register inscrapers/browser/__init__.py→SCRAPER_REGISTRY_BROWSER, thenmake deploy-browser.
Example (simple):
# scrapers/simple/example.py
from scrapers.base_scraper import BaseScraper
class ExampleScraper(BaseScraper):
def __init__(self, bq_client):
super().__init__('example', bq_client)
def run(self):
new_articles = []
# ... fetch data, use self.is_link_exists(link) for dedup ...
if new_articles:
self.save_articles(new_articles)
return self.get_stats()Register in main.py: from scrapers.simple.example import ExampleScraper and add to SCRAPER_REGISTRY.
- Cloud Functions: ~$3–5/month
- BigQuery: ~$1–2/month
- Total: about $4–7/month for the simple pipeline. Cloud Run adds cost based on usage.
- At startup, latest 20 URLs per source are loaded into memory; link existence checks use this cache to reduce BigQuery calls.
- Table is partitioned by
pub_date; all queries must include a partition filter.
bq query --use_legacy_sql=false \
"SELECT source, COUNT(*) as count
FROM \`YOUR_PROJECT_ID.news_project.news_articles\`
WHERE DATE(pub_date) >= CURRENT_DATE()
GROUP BY source"- BigQuery errors: Check table exists (
bq show news_project.news_articles) and service account permissions. - Timeouts: Increase function timeout (e.g.
--timeout=540s) or reducemax_workersinconfig.yaml.
MIT License. This is a learning project; if you believe any use infringes your rights, please contact the author for immediate removal.