A fast command-line tool that hunts and downloads files from GitHub repositories and websites. Supports PDFs, text documents, images, and videos.
Requirements: Rust 1.70+ — install via rustup.rs
git clone https://github.com/yourname/marcopolo
cd marcopolo
cargo install --path .Verify it works:
marcopolo --version
# marcopolo 0.3.0Just point marcopolo at any URL and it handles the rest:
marcopolo https://github.com/owner/repoThat's it. It crawls the page and downloads everything it finds.
Want to target specific file types? Use --type:
| Flag | What it grabs |
|---|---|
--type pdf |
PDF documents |
--type img |
Images (png, jpg, gif, webp, svg) |
--type video |
Video files (mp4, mkv, avi, mov) |
--type audio |
Audio files (mp3, wav, flac, ogg) |
--type zip |
Archives (zip, tar, gz, rar, 7z) |
--type doc |
Word / text docs (docx, doc, txt, odt) |
--type code |
Source files (rs, py, js, ts, go, cpp) |
--type data |
Data files (csv, json, xml, yaml) |
--type all |
Everything |
| Flag | Description |
|---|---|
--list |
Preview files without downloading |
--depth <n> |
How many links deep to crawl (default: 1) |
--out <dir> |
Output directory (default: ./downloads) |
Marcopolo operates in two modes depending on the URL you give it:
GitHub mode — when the URL contains github.com:
- Scans the full repository tree for committed files
- Decodes the README and extracts all hyperlinks
- Scans all GitHub Release assets
- All three run in parallel
Web mode — for any other URL:
- Checks
/sitemap.xmlat the root domain first (fast path) - BFS-crawls from the landing page up to
--depthlevels deep - Only follows same-origin links — never crawls external sites
All discovered files are deduplicated by URL, then downloaded concurrently
(4 at a time by default) into ./downloads/.
| Flag | Default | Description |
|---|---|---|
--type, -t |
pdf |
File type(s) to hunt. Repeatable. |
--out, -o |
downloads |
Output directory |
--depth |
1 |
BFS crawl depth (web mode only) |
--delay |
none | Milliseconds between downloads |
--continue |
off | Resume partially downloaded files |
--retries |
3 |
Retry attempts on failure |
--list |
off | Dry run — list files without downloading |
--filter |
none | Only download files matching this keyword |
--token |
none | GitHub personal access token |
| Flag value | Extensions |
|---|---|
pdf |
.pdf |
text |
.txt .md .epub .doc .docx .csv .rst |
img |
.jpg .jpeg .png .gif .svg .webp .bmp .ico |
video |
.mp4 .mkv .avi .mov .webm .flv .m4v |
marcopolo https://github.com/varunkashyapks/BooksDownloads all .pdf files committed in the repo into ./downloads/.
marcopolo "https://github.com/Carl-McBride-Ellis/Compendium-of-free-ML-reading-resources?tab=readme-ov-file"marcopolo decodes the README, finds all https://...pdf links, and downloads them.
marcopolo https://themlbook.com/wiki/doku.phpScrapes the page and one level of internal links for PDF hrefs.
marcopolo https://somesite.com/resources --depth 3Follows links up to 3 levels deep from the landing page.
marcopolo https://github.com/owner/repo --type imgmarcopolo https://github.com/owner/repo --type pdf --type text
marcopolo https://somesite.com --type pdf --type img --type videomarcopolo https://github.com/owner/repo --listPrints the filename and URL of every discovered file. Nothing is downloaded.
marcopolo https://github.com/owner/repo --filter "transformer"Only downloads files whose filename contains transformer (case-insensitive).
Combine with --list to preview the filtered results first:
marcopolo https://github.com/owner/repo --filter "transformer" --listmarcopolo https://github.com/owner/repo --out ~/papers
marcopolo https://somesite.com --out ./my-downloads/site-filesThe folder is created automatically if it does not exist.
Without a token, GitHub allows 60 API requests per hour. With a token, the limit raises to 5,000 per hour.
Generate one at: github.com/settings/tokens
No scopes needed for public repos. Add repo scope for private repos.
marcopolo https://github.com/owner/repo --token ghp_xxxxxxxxxxxxxxxxxxxxmarcopolo https://github.com/owner/repo --continueSends a Range header and appends bytes to partially downloaded files
instead of restarting from zero.
marcopolo https://somesite.com --delay 500Waits 500ms before each download. Useful for sites that rate-limit scrapers.
marcopolo https://github.com/owner/repo --retries 5Retries failed downloads up to 5 times with exponential back-off (500ms, 1s, 2s…).
4xx errors (404, 403) are never retried — the link is simply dead.
marcopolo "https://github.com/Carl-McBride-Ellis/Compendium-of-free-ML-reading-resources" \
--token ghp_xxx \
--out ~/ml-papers \
--retries 5 \
--delay 200marcopolo https://github.com/varunkashyapks/Books \
--type pdf \
--out ~/books \
--continuemarcopolo https://docs.someproject.org \
--type text \
--depth 2 \
--delay 300 \
--out ./docs-backupmarcopolo https://github.com/owner/design-assets \
--type img \
--filter "logo" \
--out ./logos# Step 1 — see what's there
marcopolo https://somesite.com/resources --type pdf --depth 2 --list
# Step 2 — download only what you want
marcopolo https://somesite.com/resources --type pdf --depth 2 --filter "2024"Marcopolo now includes a powerful find command to search for free books and documents across the web's largest open libraries.
When you use find, Marcopolo queries these sources in parallel:
- Internet Archive (archive.org) — Reliable JSON API
- Open Library (openlibrary.org) — Curated metadata
- Project Gutenberg (gutenberg.org) — Via Gutendex API
- Anna's Archive (annas-archive.org) — Real-time HTML scraping of the largest catalog
- Github (Github.com) — Searches PDFs with similar names on github
- Googlescholar (Googlescholar.com) — Searches PDFs / Books / Articles with similar names on Googlescholar
- Duckduckgo (Duckduckgo.com) — Searches PDFs with similar names using Duckduckgo 0% data collection
Search and preview results (default):
marcopolo find "Clean Code"Download the top match from each source:
marcopolo find "The Pragmatic Programmer" --get --out ~/my-booksRestrict search to a specific source:
marcopolo find "Computer Systems" --source archiveOptions for --source: archive, openlibrary, gutenberg, annas, github, googlescholar, duckduckgo
List results for a specific query without downloading:
marcopolo find "Rust Programming" --listMarcopolo now includes specialized sources that excel at finding PDFs across the broader internet:
Search GitHub repositories for committed PDFs:
marcopolo find "machine learning" --source github --listSearch Google Scholar for academic papers:
marcopolo find "transformer networks" --source googlescholar --listSearch DuckDuckGo specifically for PDF files:
marcopolo find "rustlang" --source duckduckgo --listThese are printed during a run and are normal. They do not crash marcopolo.
| Error | Meaning |
|---|---|
404 Not Found |
The file was moved or deleted on the remote server |
403 Forbidden |
The server blocks direct downloads |
500 Domain Not Found |
The website no longer exists |
error sending request |
The server is unreachable or timed out |
Every file that can be downloaded will be. Dead links are skipped and reported.
Check what was successfully saved:
ls -lh downloads/
ls downloads/ | wc -lAfter pulling new code from the same clone link or making changes:
cargo install --path .Use inside marcopolo's folder, this rebuilds in release mode and overwrites the global binary automatically.
