An end-to-end AI audiobook generator with a Gradio web UI and a Headless CLI pipeline. Upload any book, clone a narrator voice, clean it up, and generate a chapterized, mastered audiobook — all locally or in Google Colab / Kaggle, no cloud APIs needed.
- Fail-Safe Chapter Retry System with Exponential Backoff (
pipeline.py,progress_io.py) — Configurable per-chapter retries (max_chapter_retries: int = 2) with CUDA cache clearing (torch.cuda.empty_cache()andgc.collect()), exponential backoff delay, thread-safe retry persistence (update_chapter_retry), and an automatic end-of-run retry pass (retry_failed_at_end). - Six TTS engines behind one interface — Qwen3-TTS (default), IndexTTS-2.5, MOSS-TTS, OmniVoice, Fish Audio S2 Pro, Higgs Audio v3 (plus F5-TTS). Every engine's own controls are exposed; see TTS engines.
- Natural pacing — a paragraph's sentences are spoken together, paragraphs get their own pause, and dialogue tags stay with their quote; subtitles keep sentence-level timing.
- Chunk verification — every chunk is checked for truncation, runaway generation and silence (optionally transcribed with Whisper) and re-synthesized when it fails.
- Spoken-form text — numerals, currency, dates, roman numerals and abbreviations are rewritten the way a narrator would read them (English).
- One-file audiobooks with chapter markers — M4B/MP3 output with a navigable chapter list, cover and tags.
- Single-GPU VRAM Optimization & Dynamic Batching (
gpu_pool.py) — Configurable VRAM headroom (vram_headroom_gb: float = 2.0) to compute safe batch sizes on lower-memory GPUs (e.g. 8 GB cards). Includes preflight low-VRAM auto-detection warnings (<= 8.5 GB). - Completion Validation & Top-Level Summary Finalizer (
pipeline.py) — Enforces minimum WAV file size guard (_MINIMUM_CHAPTER_WAV_BYTES = 10_000) and outputs a top-levelgeneration_summaryblock (completed/failed counts, failed chapter numbers, ISO timestamp) togeneration_progress.jsonupon run completion. - WebSocket Keep-Alive & Session End Protocol — Background keep-alive pings (15s) and dedicated
session_endWebSocket events for cloud proxy reliability (Kaggle/Colab). - Kaggle & Cloud Pre-Flight Environment Checker (
preflight.py) — Validates the full execution environment in under 30 seconds before any model loading or generation begins. Verifies Python version, PyTorch, CUDA count, bfloat16 hardware support, transformers API (BitsAndBytesConfig), soundfile, FFmpeg, voice_ref type/existence, and Python 3.12 dict view picklability. Immediately reports errors with recommendations before spending GPU hours. - Stale Checkpoint Recovery & Voice Ref Hash Caching — Validates chunk WAV integrity (
_validate_chunk_file) to detect missing or corrupted files on disk and automatically re-synthesizes them. SHA-256 hash caching (_VOICE_REF_CACHE) prevents redundant temporary WAV writes across chunk synthesis calls. - Thread-Safe Atomic Progress I/O (
progress_io.py) — Dedicated, thread-safe file I/O layer forgeneration_progress.json. Uses atomic writes (temp file +os.replace()), module-level write lock (_WRITE_LOCK), UTF-8 BOM auto-decoding, leading garbage stripping, HTML detection, and an atomic read-inside-lock pattern to eliminate TOCTOU race conditions under concurrent GPU execution. - Config Contract Schema Versioning (
AudiobookConfig) — Versioned config contract (_CONFIG_SCHEMA_VERSION = 6) with hardenedfrom_dict()construction, backward-compatible default fallback, unknown key filtering, and human-readablefield_summary()diagnostics. - Eager Multi-GPU Pool Warmup & Self-Healing (
GPUPoolManager) — Parallel model warmup viaThreadPoolExecutorduring pool creation. Failed/OOM GPU instances are automatically detected and pruned from the active pool, allowing healthy GPUs to continue synthesizing without failing the job. - 3-Stage Overlapped Pipeline & Chunk-Level Resume —
chapter_pipeline.pyimplements a high-throughput 3-stage pipeline (Stage A: CPU text preparation, Stage B: parallel GPU synthesis workers, Stage C: streaming audio mastering and async disk I/O). Sentence audio chunks are cached incrementally in.temp_chunks/— if interrupted, generation resumes mid-chapter without re-synthesizing completed chunks. - Rust PyO3 SIMD Acceleration (
audiobook_rust) — High-performance Rust extension providing 5.5× faster audio mastering, SIMD sentence splitting, and ultra-fast text normalization compiled withmaturin(with transparent Python fallbacks). - Headless CLI generation (
cli.py) — Run audiobook generation headless in cloud environments (Kaggle/Colab) or terminal without launching or maintaining a web browser interface. Includes flags for cover art injection (--embed-cover-only), quantization, device selection, and progress file execution. - Cached Book Extraction —
generation_progress.jsoncaches fully extracted and segmented chapter text. Re-parsing large books on session resume is completely eliminated! - Export Config JSON — Configure voice and chapters in Gradio, then click '📋 Export Config JSON' to save settings and cached text into a self-contained JSON ready for CLI execution.
- Multi-format book support — EPUB, MOBI, PDF, DOCX, ODT, TXT.
- Smart chapter detection — EPUB/MOBI use a TOC-based chapter checklist; PDF/DOCX/ODT let you split by page ranges.
- Chapter selection memory — Selected chapters are saved into
generation_progress.jsonand automatically restored when you resume a session — no need to re-select every time. Uploading a book after restoring JSON settings preserves your selected chapter subset automatically. - JSON-First UI Workflow — Progress file upload placed at the very top of the app interface for immediate session restore before touching book uploads.
- AI text extraction — 5-phase pipeline (Docling + OCR + ML classification + heuristic normalization) produces clean, TTS-ready text.
- EPUB image OCR — EasyOCR reads text embedded in images inside EPUBs.
- Voice Design & Cloning — Clone from a reference WAV or prompt an entire new voice using Qwen3-TTS. Qwen3-TTS speaks 10 languages (English, Chinese, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian); other engines add more, including Hindi.
- Voice preprocessing — 7-step audio cleaning pipeline: noise reduction, noise gate, high-pass filter, silence removal, normalization, formant shifting, resampling.
- Voice test tab — Type any sentence and preview the cloned voice before generating. Includes language-labeled premium timbres for optimized results.
- Preview mode — See chapter list with character + word counts before committing to a full audiobook run.
- Pronunciation fixes — Upload a
.txtfile withsearch==replacepairs to fix how the TTS pronounces specific words. - True Multi-GPU Parallel Utilization (
GPUPoolManager) — Automatically detects all available GPUs (e.g. Kaggle T4 × 2) and assigns dedicated model instances per GPU. Enables true simultaneous work-stealing parallel execution across GPUs (with VRAM safety guards) for up to 2× speedups without OOMs. - Fast Diagnostic Pre-Run Verification — Included
scripts/colab_prerun_check.pyandscripts/kaggle_prerun_check.pyrun in <5 seconds to verify GPU allocation, VRAM limits, Rust compilation status, and GPU pool dispatch before loading heavy TTS model weights. - INT8 Quantization Support (
--quantization int8) — 8-bit model quantization viabitsandbytesreduces VRAM consumption by ~50% for memory-constrained environments. - Cover Image Embedding CLI Tool (
--embed-cover-only) — Instantly inject cover image artwork and ID3 tags into existing output audio files without running TTS generation. - Multi-Format Subtitle & Timed Lyrics Export — Generates
.lrctimed lyrics,.srtsubtitles, and.vttWebVTT files for sync with Audiobookshelf and media players. - Audiobookshelf-compatible output — Zero-padded filenames + full ID3 tags (title, author, album, track) ready to drop into Audiobookshelf.
- Mastered Output & Single File Mode — Output mastered MP3, FLAC, WAV, or M4B files. Optionally combine all chapters into a massive single unified file with one click.
- Live generation log & Decimal Progress — Stream progress in real time with a sub-chapter decimal progress bar (e.g. 74.52%) and detailed live logs.
- FastAPI / WebSocket Orchestration Server — Offloads heavy GPU jobs from Gradio to a detached FastAPI backend. Protects GPU VRAM limits via a concurrent task queue while providing real-time log and rendering progress updates via WebSockets.
torch.compile()Speed Optimization — Enable kernel fusion rendering to compile Qwen3 TTS model via the GPU compiler, speeding up audio generation throughput on RTX GPUs.- Smart Attention Backend — Automatically detects whether
flash_attnis installed. Uses Flash Attention 2 if available, otherwise gracefully falls back to PyTorch's built-in SDPA — no crashes on T4 or other GPUs that don't haveflash_attn. - Re-generate missing files control — New checkbox on the Generate tab lets you decide whether chapters marked "completed" but missing audio should be re-generated or silently skipped.
- Multi-GPU by default — every GPU pulls batches from one shared queue, so two T4s share a chapter and a failing GPU hands its work to the other. An optional
automode gives each chapter as many GPUs as it has batches to fill: short chapters run one per GPU, long ones are shared. - Books in other languages — chunk length follows the script (a Chinese character is a whole syllable, so Chinese, Japanese and Korean chunks are shorter in characters), and sentence breaks handle CJK quotes, French guillemets, dialogue dashes and the Devanagari danda.
- Google Colab & Kaggle support — Full end-to-end pipeline works directly in Google Colab (
notebooks/AudiobookMaker_Colab.ipynb) and Kaggle (notebooks/AudiobookMaker_Kaggle.ipynb) with public shareable Gradio links.
AudiobookMaker/
├── install.sh / install.bat ← One-click installer (detects OS + GPU)
├── run.sh / run.bat ← Start app + open browser automatically
├── app.py ← Gradio UI entry point
├── cli.py ← Headless CLI entry point
├── start_api.py ← FastAPI orchestration server launcher
├── requirements.txt ← Core dependencies + Qwen3-TTS
├── requirements/ ← One file per optional engine (tts-<engine>.txt)
├── notebooks/
│ ├── AudiobookMaker_Colab.ipynb ← Google Colab notebook (full pipeline, shareable link)
│ └── AudiobookMaker_Kaggle.ipynb ← Kaggle notebook (dual-GPU T4x2 support, shareable link)
├── scripts/
│ ├── colab_prerun_check.py ← Colab diagnostic check (<5s)
│ ├── kaggle_prerun_check.py ← Kaggle diagnostic check (<5s)
│ └── lrc_to_srt_converter.py ← Convert .lrc files to .srt
├── audiobook_rust/ ← Rust PyO3 extension: sentence splitting, text cleaning, mastering
├── docs/
│ └── preview/ ← UI screenshots
├── api/
│ ├── server.py ← FastAPI server (task queue, WebSocket progress streaming)
│ └── worker.py ← Background task consumer (concurrent GPU queue)
├── tests/ ← pytest suite (mock engine; no GPU needed)
└── audiobook_factory/
├── preflight.py ← Automated pre-flight environment validator (<30s check)
├── pipeline.py ← Audiobook generation orchestrator
│ (AudiobookConfig, CancelToken, run_pipeline)
├── chapter_pipeline.py ← 3-stage per-chapter pipeline & chunk-level resume
├── chunk_planner.py ← Cuts a chapter into TTS chunks (sentence packing, pauses, script-aware length)
├── chunk_verifier.py ← Checks each chunk (duration, optional Whisper transcript) and triggers re-synthesis
├── speech_text.py ← Rewrites numerals, dates, currency for narration (English)
├── loudness.py ← Loudness normalisation with a look-ahead limiter (Python twin of the Rust code)
├── extractor_engine.py ← Core text extraction engine
│ (DocumentIngestor, MLClassifier, TextNormalizer)
├── text_extractor.py ← Public API: scan() + extract()
├── voice_preprocessor.py ← Reference-voice cleaning pipeline
├── gpu_pool.py ← GPU detection, provider pool, device dispatch
├── filename_sanitizer.py ← Cross-platform, Audiobookshelf-compatible filenames
├── text_processing.py ← Sentence splitting + normalisation (Rust or NLTK)
├── ffmpeg_utils.py ← FFmpeg encoding helpers
├── progress_io.py ← Thread-safe atomic progress JSON I/O
├── utils.py ← Shared utilities (LRC timestamping, SRT formatting)
└── tts_providers/ ← One module per TTS engine
├── base_tts_provider.py ← Provider contract + get_tts_provider() factory
├── registry.py ← Which engines exist (lazy imports)
├── qwen_provider.py ← Qwen3-TTS (default)
├── indextts_provider.py ← IndexTTS-2.5 / IndexTTS-2
├── moss_provider.py ← MOSS-TTS
├── omnivoice_provider.py ← OmniVoice
├── fish_provider.py ← Fish Audio S2 Pro
├── higgs_provider.py ← Higgs Audio v3
└── f5tts_provider.py ← F5-TTS
Engine (tts_provider_name) |
Weights licence | Commercial use | Min. VRAM* | Batched | Reference transcript | Notes |
|---|---|---|---|---|---|---|
Qwen3-TTS (qwen, default) |
Apache-2.0 | yes | 6 GB | yes | optional (improves cloning) | Voice clone, preset speakers, designed voices, saved voice presets |
IndexTTS-2.5 (indextts) |
bilibili Model Use License | yes, below 100M MAU / RMB 1B revenue, with conditions | 7 GB | no | not used | Emotion controlled separately from timbre; 22 kHz |
MOSS-TTS (moss) |
Apache-2.0 | yes | 11 GB | yes | optional | Duration control, pause tags; ~13 GB download |
OmniVoice (omnivoice) |
CC-BY-NC | no | 4 GB | yes | optional | 600+ languages; style is attribute tags, not prose |
Fish Audio S2 Pro (fish) |
Fish Audio Research License | no | 12 GB | no | required | 44.1 kHz, inline emotion tags; slower than real time on a T4 |
Higgs Audio v3 (higgs) |
Boson research / non-commercial | no (creator grant with attribution) | 11 GB | yes | optional | Ported from the reference server; no official in-process path exists |
F5-TTS (f5tts) |
CC-BY-NC-4.0 | no | 3 GB | no | optional | English / Chinese |
* Estimated from weight sizes. Run python cli.py --list-providers for each engine's options and install command.
One engine per environment. The engines pin incompatible transformers versions (IndexTTS 4.52, Qwen 4.57, MOSS exactly 5.0.0, OmniVoice and Higgs 5.3+), so install only the one you use on top of requirements.txt:
pip install -r requirements/tts-<engine>.txt # indextts | moss | omnivoice | fish | higgsrequirements/tts-<engine>.txt lists any second step an engine needs. Read the licence of an engine before publishing audio made with it: the non-commercial ones do not allow selling the result.
Numbers from test runs of this code on two T4 GPUs with one 22-second English narrator clip. Speed is audio produced per second of synthesis; above 1× is faster than real time.
| Engine | Speed on 2 GPUs | Gain from the 2nd GPU | English word errors | VRAM per GPU |
|---|---|---|---|---|
| Qwen3-TTS 1.7B | 3.3× (long chapters) | 1.6× | 3% | 6–13 GB with batch size |
| Higgs Audio v3 | 3.1× | 1.6× | 4% | 9.4 GB |
| OmniVoice | 0.6× | 1.95× | 3% | 4–5 GB |
| IndexTTS-2.5 | 0.6× | 1.85× | 4% | 5.7 GB |
| MOSS-TTS | 0.6× | 1.85× | 3% | 12.6 GB |
| Fish Audio S2 Pro | 0.5× | 1.45× | 4% | 11.4 GB |
- A whole 20-page book (ten chapters, 32 minutes of audio) with Qwen3-TTS took 19 minutes, with 2% word errors over the whole book and no out-of-memory retries.
- Short chapters and two GPUs. A Qwen batch takes about the same time whether it holds 7 chunks or 17, so a chapter that fits in one batch is no faster on two GPUs. For books with short chapters set Parallelism to
auto(orchapters): short chapters then run one per GPU.autowas added after these runs and has only been exercised with simulated GPUs. - Other languages with an English narrator clip (Qwen3-TTS). Russian and Korean read well (5–6% errors). French, Japanese and Chinese were intelligible but accented, and digits were sometimes read in English. A narrator clip in the book's own language is the first thing to try.
- Hindi. By ear, Higgs Audio v3 was the only engine with correct pronunciation and no dropped words; OmniVoice, MOSS-TTS and Fish were not good enough. Qwen3-TTS and IndexTTS do not support Hindi.
- Voice similarity. A speaker-verification model scored English clones 0.99 against the narrator clip (two unrelated voices scored 0.63 and 0.83); clones in other languages scored 0.92–0.97.
- Python 3.11+
- NVIDIA GPU with 6 GB+ VRAM (strongly recommended — CPU is very slow for Qwen3-TTS)
- CUDA Toolkit 11.8+
- FFmpeg — the installer tries to handle this automatically
Note: Flash Attention 2 is optional. If the
flash_attnpackage is not installed (e.g. on T4 GPUs in Colab/Kaggle), the app automatically falls back to PyTorch SDPA — no manual action needed.
git clone https://github.com/MSpider3/AudiobookMaker.git
cd AudiobookMakerThe installer automatically:
- Detects your OS and installs Python 3.11 via the native package manager
- Creates a virtual environment
- Detects your GPU and installs the correct PyTorch (CUDA 12.1, CUDA 11.8, or CPU)
- Installs all dependencies from
requirements.txt - Detects if the Rust toolchain (cargo) is installed, compiling the high-performance PyO3 extension (
audiobook_rust) in release mode (with fallback to pure Python if Rust is not present) - Installs FFmpeg if missing
install.batchmod +x install.sh
./install.shThe run script activates the environment, starts the server, and opens your browser automatically.
run.batchmod +x run.sh
./run.shYour browser will open at http://localhost:7860 automatically.
For faster generation or execution in cloud environments (Google Colab / Kaggle notebooks) where Gradio tunnels might disconnect, you can run generation headless via cli.py:
# Straight from a book file, no JSON needed:
python cli.py --book /path/to/book.epub --voice-file /path/to/narrator.wav \
--voice-transcript-file /path/to/narrator.txt \
--provider qwen --language English --output-format m4b --single-file
# See the engines, their licences and options:
python cli.py --list-providers
# Check what a run would do without generating anything:
python cli.py --book /path/to/book.epub --voice-file narrator.wav --dry-run
# Basic run with progress JSON (uses cached chapter text):
python cli.py audiobook_output/MyBook/generation_progress.json
# Override book path (for cover image extraction) and narrator voice:
python cli.py generation_progress.json \
--book-path /path/to/book.epub \
--voice-file /path/to/voice.wav
# Override generation parameters on the fly:
python cli.py generation_progress.json \
--worker-count 4 \
--output-format mp3 \
--output-dir ./my_output
# Enable INT8 quantization (reduces VRAM by ~50%):
python cli.py generation_progress.json --quantization int8
# Embed cover image into pre-generated audio files without running TTS:
python cli.py generation_progress.json \
--cover-image /path/to/cover.jpg \
--embed-cover-only
# Force re-processing all chapters from scratch:
python cli.py generation_progress.json --force-reprocess| Flag | Description |
|---|---|
config_json |
Path to generation_progress.json (exported by the UI). Optional when --book is given. |
--book |
Generate straight from a book file (EPUB, PDF, TXT, DOCX, ODT, MOBI). |
--book-path |
Book file to use with a JSON (for the cover, or when the JSON has no chapter text). |
--chapters |
Only run these chapter numbers, e.g. 1-12 or 1,3,5. |
--redo |
Regenerate these chapters even if they are finished. |
--provider / --tts-model-name / --tts-option KEY=VALUE |
Engine, model and engine-specific options (see --list-providers). |
--language |
Language of the text, e.g. English, Hindi. |
--voice-file |
Narrator reference clip to clone. |
--voice-transcript / --voice-transcript-file |
What is said in the reference clip, as text or as a file. |
--voice-preset |
Saved voice preset of the engine, used instead of a clip. |
--single-file |
Combine all chapters into one file with chapter markers. |
--speed, --pause, --para-pause, --max-len |
Narration pacing and chunk length. |
--verify |
Check each chunk: off, duration (default) or asr (Whisper transcript). |
--gpu-count, --batch-size |
GPUs to use and chunks per batch (0 = automatic). |
--parallel-mode |
chunks (all GPUs share one chapter, default), chapters (one chapter per GPU) or auto (short chapters one per GPU, long ones shared). |
--dry-run / --list-providers |
Print the resolved configuration / the available engines, then exit. |
--local |
Run in this process even when the API server is up. |
--output-dir |
Override destination output directory. |
--output-format |
Override audio output format (mp3, flac, wav, m4b). |
--worker-count |
Override parallel worker count. |
--device |
Select compute device (cuda or cpu). |
--quantization |
Model quantization mode (none or int8). |
--cover-image |
Supply or override cover image file (.jpg, .png, .webp). |
--embed-cover-only |
Instantly embed cover art into existing output audio without TTS. |
--no-resume-chunks |
Disable chunk-level disk cache resume and re-synthesize all chunks. |
--force-reprocess |
Force re-extraction and re-synthesis of all chapters. |
- Configure your settings, narrator voice, and chapter selections in the Gradio Web UI.
- In the Generate tab, click 📋 Export Config JSON. This saves
generation_progress.jsoncontaining all settings and embedded chapter text. - Close or stop the Gradio app.
- Run
python cli.py generation_progress.jsonin your terminal or cloud notebook.
You can run AudiobookMaker entirely in cloud notebooks — no local GPU or installation required.
- Open
notebooks/AudiobookMaker_Colab.ipynbin Google Colab. - Enable a GPU runtime: Go to Runtime → Change runtime type → T4 GPU → Save.
- Run pre-run diagnostics (optional but recommended):
!python scripts/colab_prerun_check.py
- Run all notebook cells in order. The notebook will automatically install dependencies, compile the Rust PyO3 extension, launch the background FastAPI server, and provide a public
gradio.livelink.
- Upload or open
notebooks/AudiobookMaker_Kaggle.ipynbin Kaggle. - Configure Session Options (Right Sidebar):
- Accelerator: Set to GPU T4x2 (enables dual-GPU parallel synthesis).
- Internet: Toggle ON (required to clone repo and download models).
- Run pre-run diagnostics:
Verifies both T4 GPUs, VRAM metrics, PyO3 bindings, and multi-GPU pool dispatch in <5 seconds.
!python scripts/kaggle_prerun_check.py
- Run all notebook cells to start the FastAPI server and launch the Gradio Web UI link.
| Feature / Behavior | Detail |
|---|---|
| Multi-GPU Parallelism | Kaggle T4x2 assigns one model instance per GPU. Long chapters gain about 1.6× from the second GPU; for books with short chapters choose Parallelism auto. |
| Other engines | Each optional engine needs its own transformers version, so install one per session (the notebooks have a cell for it). |
| Flash Attention | T4 GPUs don't have flash_attn pre-installed. The app auto-detects this and uses PyTorch SDPA — generation still works flawlessly. |
| Rust Acceleration | Pre-run check compiles audiobook_rust via maturin. If compilation fails, pure-Python fallbacks activate automatically. |
| Chunk-Level Resume | Progress and chunk audio (.temp_chunks/) are cached to disk so you can resume interrupted cloud sessions without losing progress. |
- Upload your book file (
.epub,.mobi,.pdf,.docx,.odt,.txt) - EPUB / MOBI with TOC → A chapter checklist appears. Tick the chapters you want to convert. Use Select All / Deselect All for quick bulk selection.
- PDF / DOCX / ODT / MOBI (no TOC) → Enter page ranges, e.g.
1-50, 51-120, 121-250. Each range becomes a separate chapter file. - TXT → No page structure; the whole file becomes one audio file automatically.
- Language Selection → Choose the language of your book from the dropdown. This tells the TTS engine which phonetic dictionary to use.
- Fill in book title, author, choose output format and LUFS loudness target.
Tip: Your chapter selection is automatically saved to
generation_progress.jsonwhen generation starts. Upload that file later to restore your exact chapter picks without re-selecting.
Upload your raw voice WAV and run any combination of these steps:
| Step | What it does |
|---|---|
| Noise Reduction | Reduces background hiss/hum |
| Noise Gate | Silences frames below a dB threshold |
| High-Pass Filter | Removes low-frequency rumble |
| Silence Removal | Strips long silences between words |
| Normalize Volume | Peaks at your chosen dBFS |
| Formant Shift | Adjust voice gender/timbre (experimental) |
| Resample | Convert to 22k / 44.1k / 48k Hz |
Click ▶ Preview Processed Audio to hear the result, then 💾 Use as narrator voice to pass it to the next tab.
- Upload or carry over the processed voice WAV.
- Select a TTS Model Variant (Base for cloning, CustomVoice/VoiceDesign for prompting).
- Choose a Premium Timbre (if using CustomVoice). Choices are prefixed with their native language (e.g.,
[English] ryan,[Japanese] ono_anna) for the best quality match. - Adjust TTS tuning parameters (speed, temperature, top-p, sentence/paragraph pauses).
- Type any sentence in the Voice Test box and click ▶ Test Voice to hear a preview.
- Max chunk length — TTS input character limit per sentence chunk (default 399).
- Parallel chapter workers — Process 1–8 workers simultaneously. Automatically defaults to
min(gpu_count * 4, 8)based on detected GPUs. Thanks to our Multi-GPU Pool architecture (GPUPoolManager), workers dynamically stream work across all available GPUs in parallel with optimal VRAM management. - TTS Provider — Qwen3-TTS by default; IndexTTS-2.5, MOSS-TTS, OmniVoice, Fish Audio S2 Pro, Higgs Audio v3 and F5-TTS once installed (see TTS engines).
- Parallelism — How several GPUs are used.
chunks: every GPU shares one chapter (default).chapters: one chapter per GPU.auto: short chapters one per GPU, long ones shared. - EasyOCR — Enable to extract text from images embedded inside EPUB files.
- Force reprocess — Re-extract text even if cached output exists.
- Export chapter text / Subtitles — Option to write
.txtchapter text,.lrctimed lyrics,.srtsubtitles, and.vttWebVTT subtitles alongside audio. - Pronunciation fix file — Upload a
.txtwith one fix per line insearch==replaceformat (regex supported). Comments start with#.# Fix common TTS mispronunciations Barbadoes==Barbayduss N\.E\.==north east Dr\.==Doctor - Resume / Sync Progress — Upload an existing
generation_progress.jsonto resume a previous session. All settings (voice, model, format, chapter selection) are automatically restored to the UI.
- Click 🔍 Preview Chapters to see a table of chapter titles, character counts, word counts, and sentence counts — without generating any audio. Great for checking your chapter selections.
- Click 🎧 Generate Audiobook to start the full pipeline.
- Watch the Live Decimal Progress Bar and log stream.
- Use ⛔ Cancel to stop at any time.
- 🔄 Re-generate completed chapters whose audio file is missing (new checkbox):
- Checked (default): If a chapter is marked
completedin the progress JSON but the audio file is missing on disk, it will be automatically re-generated. - Unchecked: Chapters marked completed but with missing files are silently skipped — useful when files exist in a different location.
- Checked (default): If a chapter is marked
- When complete, download individual chapter files or use ⬇ Download All (ZIP)
Edit audiobook_factory/tts_providers/qwen_provider.py:
# In _load_base_model():
"Qwen/Qwen3-TTS-12Hz-1.7B-Base" # replace with any compatible Qwen3 checkpoint
# In _run_genesis():
"Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign" # design model used once for voice genesis- Create
audiobook_factory/tts_providers/my_provider.py - Subclass
BaseTTSProviderand implementsynthesize(),estimate_cost(),get_name() - Register the name in
base_tts_provider.get_tts_provider() - Add the name to the
tts_provider_dddropdown inapp.py
Edit audiobook_factory/extractor_engine.py:
TextNormalizer._strip_noise()— add/remove markdown patterns to cleanTextNormalizer._fix_isolated_capitals()— font-kerning fixes (e.g.T HE→THE)_SKIP_TOC_TITLEregex — controls which TOC entries are excluded (copyright, gallery, etc.)MLClassifier.predict_is_chapter()— swap in a trained XGBoost model here when ready
Edit audiobook_factory/pipeline.py:
lufs: int = -18 # loudness target
true_peak: float = -1.5 # max true peak dBTPOr adjust these in the UI (LUFS slider in Book tab, True Peak in Advanced tab).
When a chapter cannot reach the loudness target under the true-peak ceiling with gain alone, the peaks in the way are turned down by a look-ahead limiter (audiobook_factory/loudness.py, mirrored in audiobook_rust/src/audio/master.rs).
Edit audiobook_factory/ffmpeg_utils.py — add a new entry to get_format_settings().
Edit audiobook_factory/voice_preprocessor.py:
- Each step is a standalone function — easy to add, remove, or reorder
PreprocessConfigdataclass controls all defaults
| Format | Chapter Detection | Fallback |
|---|---|---|
| EPUB | ✅ TOC chapter list | — |
| MOBI | ✅ Try TOC | Page-range picker |
| ❌ | Page-range picker | |
| DOCX | ❌ | Page-range picker |
| ODT | ❌ | Page-range picker |
| TXT | ❌ | Whole book |
| Format | Notes |
|---|---|
| MP3 | Default, most compatible |
| FLAC | Lossless |
| WAV | Uncompressed |
| M4B | Audiobook format with chapter markers (Apple Books) |
Audiobookshelf is a self-hosted audiobook library server. AudiobookMaker generates output that Audiobookshelf automatically detects:
- Drop the output folder into your Audiobookshelf library directory
- Audiobookshelf will auto-scan and import it as a book
- Each chapter file has the correct ID3 metadata (title, author, album, track number) so chapter ordering and library display work correctly out of the box
Output filenames follow the {NNNN}_{Chapter_Title}.mp3 format Audiobookshelf expects.
Modern versions of Gradio implement sandbox security checks that restrict browsers from loading server-generated files directly. To ensure seamless operation, AudiobookMaker automatically whitelists the project root directory using allowed_paths=[_ROOT] inside app.py. This enables:
- Transferring processed audio from the Voice Preprocessing tab directly to the Voice Studio tab without errors.
- Viewing and downloading final generated output audio/ZIP chapter packages directly from the web interface.
- Engines: IndexTTS-2.5, MOSS-TTS, OmniVoice, Fish Audio S2 Pro and Higgs Audio v3 join Qwen3-TTS behind one provider contract; Qwen3-TTS gained preset speakers, designed voices and saved voice presets. VibeVoice was removed.
- Narration: sentences of a paragraph are spoken together, paragraphs get their own pause, numerals and dates are rewritten for speech, and every chunk is checked (and re-synthesized if it is cut short, silent or runs away).
- Audio: chapters now reach the loudness target through a look-ahead limiter; single-file M4B/MP3 with chapter markers; speed control for every engine.
- Multi-GPU: one shared queue per chapter, even sharing on short chapters, and an optional
automode that hands short chapters out one per GPU. - Books: chapter detection for TXT, DOCX, ODT and PDF, real MOBI support, and correct chunking and sentence breaks for Chinese, Japanese, Korean, Hindi, French and Russian.
- Tools: the CLI runs straight from a book file (
--book), the API lists engines and serves result files, and the repository layout was tidied (notebooks/,scripts/). - Licence: relicensed from Apache-2.0 to AGPL-3.0-or-later.
See CHANGELOG.md for the full list.
- Fail-Safe Chapter Retry System (
pipeline.py,progress_io.py): Automatic chapter retries (max_chapter_retries: int = 2) with CUDA cache clearing (torch.cuda.empty_cache()andgc.collect()), exponential backoff delay, retry persistence (update_chapter_retry), and end-of-run retry pass (retry_failed_at_end). - Pluggable Provider Architecture & New TTS Providers: Native support for VibeVoice-1.5B (
bezzam/VibeVoice-1.5B-hf) and F5-TTS (SWAVE-LAB/F5-TTS) zero-shot voice cloning alongside Qwen3-TTS. - Single-GPU VRAM Optimization: Configurable reserved VRAM headroom (
vram_headroom_gb: float = 2.0) to calculate safe batch sizes on lower-memory GPUs (e.g. 8 GB cards). Low-VRAM preflight warnings (<= 8.5 GB). - Completion Validation & Progress Finalizer: Enforces minimum WAV size guard (
_MINIMUM_CHAPTER_WAV_BYTES = 10_000) and outputs a top-levelgeneration_summaryblock (completed/failed counts, failed chapter numbers, ISO timestamp) togeneration_progress.jsonupon run completion. - WebSocket Keep-Alive & Session End Protocol: Background 15s keep-alive pings (
{"type": "ping"}) and dedicatedsession_endWebSocket events for cloud proxy reliability (Kaggle/Colab). - FastAPI Graceful Shutdown & Task Error Isolation: Added
@app.on_event("shutdown")in FastAPI server to flag task cancellation, flush checkpoints, and shut downGPUPoolManager. Isolated worker task execution in_run_task_safely().
- Pre-Flight Environment Validator (
preflight.py):run_preflight_checks()validates Python, PyTorch, CUDA count, bfloat16 hardware support, transformers API, soundfile, FFmpeg, voice reference, and Python 3.12 dict view picklability in under 30 seconds before model loading. - Voice Reference Safety & Hash Caching: Implemented
BaseTTSProvider._validate_voice_ref()boundary validation and SHA-256 hash caching (_VOICE_REF_CACHE) in_resolve_voice_ref()to prevent redundant temp file writes on disk. - Stale Checkpoint Detection: Added
_validate_chunk_file()(minimum size guard = 1,000 bytes). Missing or corrupted chunk files on disk are automatically detected, reported with a stale checkpoint warning, and re-synthesized. - Python 3.12 Pickling Support: Extended
_sanitize_dict_keys()to coverdict_keys,dict_values, anddict_itemsview objects across model and generation configs.
- 3-Stage Overlapped Architecture: Implemented
chapter_pipeline.py, featuring CPU text preparation (Stage A), parallel GPU batch synthesis workers (Stage B), and streaming partial mastering with async disk I/O (Stage C). - Chunk-Level Mid-Chapter Resumption: Synthesized audio sentence chunks are cached to
.temp_chunks/during generation. Interrupted runs seamlessly pick up from the exact sentence without re-synthesizing completed chunks.
- PyO3 Extension Module: SIMD-accelerated sentence splitting, text cleaning, text normalization, and 5.5× faster audio mastering written in Rust.
- Transparent Fallback: Graceful fallback to pure-Python implementations if Rust compilation is omitted.
- Kaggle Notebook (
notebooks/AudiobookMaker_Kaggle.ipynb): Native Kaggle setup supporting GPU T4x2 dual-GPU parallel execution. - Fast Diagnostic Checkers: Added
scripts/colab_prerun_check.pyandscripts/kaggle_prerun_check.pyto test hardware, PyO3 bindings, and GPU pool dispatch in under 5 seconds before loading model weights.
- Cover Embedding Tool:
--embed-cover-onlyflag allows instant injection of cover art and ID3 metadata into existing audio files without triggering model load. - INT8 Quantization: Added
--quantization int8flag viabitsandbytes, cutting model VRAM usage by ~50%.
- Added
.srtsubtitle and.vttWebVTT export alongside.lrctimed lyrics for maximum compatibility across media players and Audiobookshelf.
- True Multi-GPU Support: Implemented
GPUPoolManagerandProviderPool(audiobook_factory/gpu_pool.py), which automatically detects all CUDA devices (e.g. dual Tesla T4s on Kaggle) and loads dedicated model instances per GPU. - Work-Stealing Task Dispatch: Chapter and chunk synthesis dynamically acquires and releases GPU provider instances from a thread-safe pool, delivering up to 2× faster synthesis on multi-GPU systems.
- Concurrent API Worker Queue: Converted
api/worker.pybackground consumer loop to run tasks concurrently up to the number of detected GPUs (asyncio.Semaphore). - Gradio Multi-GPU Status Badge: Header banner displays real-time GPU hardware detection (
GPU: cuda:0 + cuda:1 (2× parallel)), and the Advanced tab worker slider automatically defaults tomin(gpu_count * 4, 8). - API Health Endpoint Reporting:
GET /api/v1/healthnow returns detailed multi-GPU pool status and free/total VRAM metrics per device. - Warmup Threading: Server startup launches a non-blocking background thread to warm up GPU models before the first user request.
- Chapter selection is now persisted in
generation_progress.json. When you upload a progress file to resume generation, the chapter checkbox list is automatically restored to the exact same selection — no need to manually re-select chapters each time. - New "Re-generate missing files" checkbox on the Generate tab gives you control over what happens when a chapter is marked
completedbut its audio file is missing on disk.
- Flash Attention 2 auto-detection: The TTS model loader no longer requires
flash_attnto be pre-installed. It detects availability at runtime and gracefully falls back to PyTorch SDPA — preventing crashes on T4 GPUs in Colab/Kaggle and other environments. - NLTK
punkt_tabauto-download: NLTK 3.9+ requires a newpunkt_tabresource for sentence tokenization. AudiobookMaker now automatically checks for and downloads bothpunktandpunkt_tabon startup, preventing pipeline crashes on fresh environments. CancelTokenattribute fix: Resolved anAttributeErrorin the API worker that caused the cancellation flow to crash ('CancelToken' object has no attribute 'cancelled'→ fixed to use.is_cancelled).- Rust module graceful fallback: Added
hasattr()guards around all Rust extension calls (clean_text,normalize_text,split_sentences,master_audio). The pipeline continues with pure-Python implementations if the Rust module was compiled without certain functions.
This project would not have been possible without the incredible work from these projects:
Qwen3-TTS by QwenLM
The voice cloning and TTS engine powering high-quality audio generation. State-of-the-art text-to-speech with zero-shot voice cloning from a short reference clip.
Optional TTS engines; each is used under its own licence (see TTS engines).
F5-TTS by SWAVE-LAB
Fast, lightweight zero-shot text-to-speech voice cloning model integrated as a provider option in AudiobookMaker.
Mangio-RVC-Fork by Mangio621
The voice preprocessing pipeline in this project (noise reduction, noise gate, high-pass filter, silence removal, formant shifting) is directly inspired by the preprocessing architecture used in Mangio-RVC-Fork.
AGPL-3.0-or-later — see LICENSE and NOTICE.
- You may use, modify and redistribute AudiobookMaker, including commercially, as long as derived versions stay under the same licence and come with their source.
- AudiobookMaker has a web UI and an HTTP API: if you run a modified version for others over a network, you must offer them its source (AGPL section 13).
- The audiobooks you generate are yours; the licence of the software does not apply to them. The licence of the TTS engine you used does — see TTS engines.
- The optional TTS engines are installed separately under their own licences;
NOTICEgrants an additional permission to combine AudiobookMaker with them. - Releases up to v1.5.0 were published under Apache-2.0 and remain available under it.




