A WPF (C#/.NET 8) system-wide dictation tool for Windows. Hold a global hotkey, speak, and the transcript is typed or pasted into whichever application has focus. Works with multiple cloud ASR backends (Mistral Voxtral, OpenAI Whisper, ElevenLabs Scribe, Cohere Transcribe, Deepgram Nova-3, Google Chirp 3, Soniox, Modulate Velma 2, Smallest.ai Waves, Reson8) and optional local inference paths (GGUF via CrispASR/llama.cpp, incl. Qwen3-ASR 1.7B).
- Repository: https://github.com/praxeo/whisperinc
- Platform: Windows 10/11, .NET 8.0
- License: see
LICENSE
Three commands from a fresh clone on a Windows box with the .NET 8 SDK:
git clone https://github.com/praxeo/whisperinc.git
cd whisperinc
.\scripts\install.ps1 -DesktopThat publishes a self-contained build to _publish\ and creates Start Menu + Desktop shortcuts. Launch from the Start Menu. Right-click the tray icon for the support menu; hold Ctrl+Space in any window to dictate.
Uninstall with .\scripts\uninstall.ps1. It removes shortcuts and the auto-start entry but preserves %APPDATA%\.WhisperInk\ (config, history, models).
- Features
- Setting up on a new computer
- Configuring a provider
- Hotkeys & modes
- Optional local backends
- Architecture
- Configuration & data locations
- Build & publish
- Getting help
- Troubleshooting
- Design docs
- Global hotkey dictation β
Ctrl+Spaceto record and transcribe into any foreground app. - Multi-provider β cloud (Mistral, OpenAI, ElevenLabs, Cohere, Deepgram, Google Chirp 3, Soniox, Modulate, Smallest.ai, Reson8) and local (GGUF via CrispASR subprocess/server, incl. Qwen3-ASR 1.7B). Per-provider auth header, endpoint, model, temperature, and context-bias configuration.
- Batch dictation β record β POST to the active provider β paste via clipboard; works with every provider.
- Never loses a dictation β every take's audio is saved before it is sent and deleted once its text is delivered. A timeout, an outage or a crash leaves it in β» Unsent dictations (tray / right-click menu) to retry β with the usual provider, or on a local model when the cloud is down, clearly flagged as a fallback. Deadlines scale with the recording, so long takes don't time out just for being long.
- Loud failures β a distinct tone for each outcome: success, error, and a two-pulse warn for "delivered, but check it" (transcript may be incomplete, mic cut out mid-take, or the window you dictated into was no longer in front so the text was left on the clipboard instead of being pasted somewhere else).
- Context biasing β one shared term list, routed to each provider's native mechanism automatically (prompt glossary for OpenAI,
context_biasfor Mistral,keytermsfor ElevenLabs,hotwordsfor local CrispASR/Parakeet, phrase sets for Google, context terms for Soniox,custom_termsfor Modulate,phrasesfor Reson8). Providers with no biasing surface at all β Cohere Transcribe, Smallest.ai β log the ignored terms rather than dropping them silently. - History log β every transcription recorded locally, viewable from the tray.
This is the fast path if you just want to use cloud providers β no GPU, no Python, no model downloads.
- .NET 8 Desktop Runtime (x64) β https://dotnet.microsoft.com/download/dotnet/8.0 Pick "Desktop Runtime", not just the runtime, or the WPF app won't launch.
- Git for Windows (only if building from source) β https://git-scm.com/download/win
- .NET 8 SDK (only if building from source) β same download page as above.
Verify:
dotnet --list-runtimes
# should include: Microsoft.WindowsDesktop.App 8.0.xOption A β build from source (recommended while there are no published releases):
git clone https://github.com/praxeo/whisperinc.git
cd whisperinc
dotnet publish -c Release -r win-x64 --self-contained false -o publishThe binary lands in publish\WhisperInk.exe. Move that folder wherever you want the app to live (e.g. C:\Tools\WhisperInk\).
Option B β self-contained build (bundles the .NET runtime, ~80MB, skips prerequisite #1):
dotnet publish -c Release -r win-x64 --self-contained true -o publishDouble-click WhisperInk.exe. The main window opens and a keyboard hook is installed. Leave the window open (minimize it) β closing it exits the hook. Optional: add a shortcut to WhisperInk.exe in shell:startup to auto-launch at login.
Windows β Settings β Privacy & Security β Microphone β allow desktop apps. On first recording, Windows may also prompt for access.
That's the minimum. From here, configure at least one provider and you can dictate.
Right-click the window (or use the menu) and open Providersβ¦.
WhisperInk ships with several defaults populated. For the easiest path, pick one cloud provider and paste its API key:
| Provider | Where to get a key | Notes |
|---|---|---|
| Mistral | https://console.mistral.ai | Voxtral transcription. |
| OpenAI | https://platform.openai.com/api-keys | Whisper-1. Rock-solid; supports whisper_prompt vocabulary biasing. |
| ElevenLabs Scribe v2 | https://elevenlabs.io | Custom xi-api-key header; supports keyterms list (~20% cost surcharge). |
| Cohere Transcribe | https://dashboard.cohere.com | Cohere v2 endpoint; temp 0.1 default. No native vocabulary biasing. |
| Modulate Velma 2 | https://platform.modulate.ai | Three presets, one per batch model. Pick Multilingual for custom_terms vocabulary biasing, or English Fast for the lowest latency (no biasing). |
| Smallest.ai Waves | https://smallest.ai | Two presets. Pulse Pro is English-only and the better dictation model; Pulse covers 46 languages. Neither supports vocabulary biasing of any kind, so the shared Context Bias list is inert here. Note the API retains request content by default. |
| Reson8 | https://console.reson8.dev | One preset (the API has no model parameter). Real vocabulary biasing via phrases (β€250 terms), plus custom_model_id for a persistent 50,000-phrase vocabulary. Ten languages only (de/en/es/fr/fy/it/nl/pl/pt/sv) β picking any other from the Language dropdown falls back to auto-detect rather than failing. Wired but not yet live-tested. |
Then:
- Select the provider in the list β paste your API key β Save.
- Set it as Active.
- Pick your microphone from the device dropdown.
- Hold Ctrl+Space, speak, release. Text should paste into the focused window.
The debug log at %APPDATA%\.WhisperInk\debug.log is your first stop if something isn't working. It starts fresh at each launch; the previous session's is kept as debug.previous.log.
| Hotkey | Behaviour |
|---|---|
| Hold Ctrl+Space | Record for as long as held; on release, transcribe and insert. The transcript is pasted via Ctrl+V. |
- Records in memory (a copy of the latest take goes to
~/Documents/MyRecordings/temp_audio.wavfor replay/debugging). - Journals the take to
%APPDATA%\.WhisperInk\unsent\before sending it; delivered takes are deleted, failed ones kept for β» Unsent dictations. - POSTs the audio to the provider, with a deadline scaled to the recording (cloud: 20 s + 20 s per minute of audio; local: 180 s + the audio's length).
- Pastes via clipboard + simulated
Ctrl+V, with a leading space prepended β but only into the window the dictation started in, and only if it is still in front. Otherwise the text is left on the clipboard and the warn tone plays.
Everything below is optional. If you're happy with a cloud provider, skip this section.
A local FastAPI server for running Cohere Transcribe on your own GPU.
Requirements:
- Python 3.10+
- NVIDIA GPU with recent CUDA drivers (RTX 30-series or better recommended)
- 16 GB+ system RAM
cd server
python -m venv venv
venv\Scripts\activate
pip install transformers>=4.52.0 torch soundfile fastapi uvicorn
python cohere_server.py
# listens on http://127.0.0.1:8101Configure WhisperInk β add a provider:
- Base URL:
http://127.0.0.1:8101 - Transcription Endpoint:
http://127.0.0.1:8101/v1/audio/transcriptions - API Key: (blank)
- Model:
cohere-transcribe-03-2026
CrispASR is a whisper.cpp/llama.cpp fork that can run the Cohere Transcribe GGUF locally. Four provider variants ship for different deployments β CPU subprocess, HTTP server (CPU), HTTP server (CUDA), HTTP server (CUDA Q8). The CPU HTTP server (cohere-gguf-server) is the most useful default on a machine without an NVIDIA GPU.
Performance expectation: on a typical laptop CPU (8 threads, no GPU), the Q5_0 build runs at roughly real-time β a few-second dictation burst transcribes in a couple of seconds. Cloud is still snappier for long clips, but for normal dictation bursts local CPU is perfectly usable; the tradeoff is latency vs. offline/privacy, not "fast vs. unusable."
- Visual Studio 2022 with the "Desktop development with C++" workload (or standalone Build Tools 2022).
- CMake 3.14+ β https://cmake.org/download/
- Git
- ~2 GB disk for the build + model.
For CUDA variants, also: CUDA Toolkit 12.x and an NVIDIA GPU with recent drivers.
cd path\to\whisperinc
.\scripts\download-cohere-gguf.ps1This fetches cohere-transcribe-q5_0.gguf (~1.45 GB) from HuggingFace (cstr/cohere-transcribe-03-2026-GGUF) into %APPDATA%\.WhisperInk\cohere-gguf\. Q5_0 is the sweet spot; edit $variant in the script to use q4_k (smaller) or q6_k/q8_0 (more accuracy).
.\scripts\build-crispasr.ps1What this does:
- Clones
https://github.com/CrispStrobe/CrispASRas a sibling of this repo (e.g., if whisperinc is atDocuments\GitHub\whisperinc\, CrispASR lands atDocuments\GitHub\CrispASR\). The script derives paths from$PSCommandPath, no hardcoded user paths. - Runs
cmake -B build -G "Visual Studio 17 2022" -A x64 -DGGML_CUDA=OFF -DWHISPER_BUILD_TESTS=OFFthen builds thewhisper-clitarget in Release. That target producescrispasr.exevia CMake'sOUTPUT_NAMEβ the historicalwhisper-cliname is kept only for internal linking rules. - Copies
crispasr.exeand every*.dllfrombuild\bin\Release\into%APPDATA%\.WhisperInk\cohere-gguf\β the 13-ish DLLs includeparakeet.dll,cohere.dll,crispasr.dll, theggml*.dllfamily,whisper.dll, etc. Copying onlyggml*is not enough; the binary dynamically loads the per-backend DLLs at startup and silently exits withSTATUS_DLL_NOT_FOUNDif any are missing.
For a CUDA build on a GPU box, flip -DGGML_CUDA=OFF to ON in the script. Requires CUDA Toolkit 12.x on PATH.
The -DWHISPER_BUILD_TESTS=OFF flag skips a Catch2 FetchContent step that tries to clone from GitHub during configure β saves a few hundred MB and avoids failures on firewalled networks.
After step 2, %APPDATA%\.WhisperInk\cohere-gguf\ should contain:
crispasr.exe
cohere-transcribe-q5_0.gguf (or whichever GGUF you downloaded)
crispasr.dll, whisper.dll,
ggml.dll, ggml-base.dll, ggml-cpu.dll, ggml-cuda.dll,
cublas64_12.dll, cublasLt64_12.dll, cudart64_12.dll (CUDA asset only)
In the WhisperInk UI, pick one of:
Cohere Local (CrispASR GGUF)β usesCohereGgufTranscriber, one-shot subprocess per recording (simplest, highest per-call latency).Cohere Local (CrispASR server)β usesCohereGgufServerTranscriber, lazy-startscrispasr.exe --server --host 127.0.0.1 --port 8766 -m <model> --backend cohere -l en -t 8 -npon first use and keeps it alive. Recommended for the CPU path β cuts latency by keeping the model loaded.Cohere Local (CrispASR CUDA)/Cohere Local (CrispASR CUDA Q8)β same idea but targeting a CUDA-built binary and, for Q8, a different model file.
Set it Active, mode = Batch, then Ctrl+Space to test. First call takes a few extra seconds while the server boots and loads the model; subsequent calls are just inference.
Source files worth reading if you want to customize ports, flags, or the model filename: CohereGgufTranscriber.cs, CohereGgufServerTranscriber.cs, CohereGgufCudaServerTranscriber.cs, CohereGgufCudaQ8ServerTranscriber.cs.
WhisperInk ships with a built-in Parakeet Local (CrispASR) provider that auto-spawns a crispasr.exe --server subprocess the first time you dictate with it selected, keeps the model resident between calls, and tears it down when WhisperInk exits. You don't need to keep a terminal open.
Parakeet TDT 0.6B at Q4_K is ~467 MB and noticeably faster than the 2B Cohere Transcribe model on CPU, so it's a good default for laptops without a GPU. The server exposes CrispASR's OpenAI-compatible /v1/audio/transcriptions endpoint, so WhisperInk uses its regular HTTP batch path β no realtime streaming, no special protocol.
Performance you should expect:
- CPU (modern laptop, 8 threads): RTFx 2β3Γ β a 2-second burst transcribes in <1s, a 10-second utterance in ~3β4s. Cold start adds ~2β3s model-load tax, once per session.
- CUDA (RTX 3090-class): RTFx 20β50Γ β near-instant for any reasonable dictation length. Requires a CUDA-built
crispasr.exeand a CUDA-variant Parakeet provider (not wired up by default β see "Canary / Qwen3-ASR / Voxtral" below for the cloning pattern). - Low-power ultrabook CPU (e.g., Ryzen U-series, Intel T-suffix): RTFx 1.5β2Γ. Still usable for dictation bursts; longer utterances feel slow.
You only need two files in place: crispasr.exe and a Parakeet GGUF, both under %APPDATA%\.WhisperInk\cohere-gguf\.
-
crispasr.exe β build CrispASR's current
mainbranch from https://github.com/CrispStrobe/CrispASR (it now ships the unified multi-backend binary). If you already followed the Cohere GGUF section above, you've got this. -
Parakeet GGUF β one PowerShell line:
$dir = "$env:APPDATA\.WhisperInk\cohere-gguf" New-Item -ItemType Directory -Force -Path $dir | Out-Null curl.exe -L --fail --progress-bar ` -o "$dir\parakeet-tdt-0.6b-v3-q4_k.gguf" ` "https://huggingface.co/cstr/parakeet-tdt-0.6b-v3-GGUF/resolve/main/parakeet-tdt-0.6b-v3-q4_k.gguf"
Any
parakeet-*.gguffilename in that folder will be picked up; pick a different quant (q8_0,q5_0,q4_k,f16) by changing the URL filename.
In WhisperInk β Providersβ¦ β pick Parakeet Local (CrispASR) β set it Active. Mode = Batch. Hold Ctrl+Space to dictate. First call takes a few extra seconds while the server boots and loads the model; subsequent calls are just inference.
Changing the port in the provider's Base URL field (default http://localhost:8103) is honored β the spawned server binds to whatever port the UI says.
Retired 2026-06-12. Removed from the shipped defaults β
cohere-local-q6kcovers the GPU path at near-F16 accuracy,cohere-gguf-serverstays the CPU fallback. The id still loads for configs that already carry it and port 8104 stays reserved, but a fresh install will not show it. Kept here because the auto-spawn pattern it describes is exactly how every local preset works.
The second built-in auto-spawn provider. Same pattern as Parakeet β lazy-starts crispasr.exe --server on port 8104 when selected, keeps the model resident, tears it down on app exit β but runs the Cohere Transcribe 2B GGUF at Q4_K quantization (~1.4 GB). Trades off size/latency for Cohere's higher English accuracy ceiling.
Performance you should expect:
- CPU (modern laptop, 8 threads): RTFx 1β1.5Γ β roughly half Parakeet's speed at the same quant level, since the underlying model is ~3Γ larger. A 2-second burst transcribes in ~1β2s, a 10-second utterance in ~8β10s. Cold start adds ~3β5s.
- CUDA (RTX 3090-class): RTFx 15β30Γ β comfortably fast for any dictation length, once wired into a CUDA-built
crispasr.exe.
Unlike Parakeet, Cohere GGUFs don't expose backend metadata that CrispASR's auto-detect reads, so WhisperInk's spawn call passes --backend cohere explicitly. That's handled in MainWindow.xaml.cs via the backendHint parameter on CrispAsrServerTranscriber.
Same crispasr.exe as the Parakeet path. The only extra file is the Q4_K GGUF:
.\scripts\download-cohere-q4.ps1Or inline:
$dir = "$env:APPDATA\.WhisperInk\cohere-gguf"
New-Item -ItemType Directory -Force -Path $dir | Out-Null
curl.exe -L --fail --progress-bar `
-o "$dir\cohere-transcribe-q4_k.gguf" `
"https://huggingface.co/cstr/cohere-transcribe-03-2026-GGUF/resolve/main/cohere-transcribe-q4_k.gguf"The file is literally named cohere-transcribe-q4_k.gguf β the dispatch in WhisperInk looks for that exact filename, not a glob, so don't rename it.
In WhisperInk β Providersβ¦ β pick Cohere Local Q4 (CrispASR) β Active, Batch. Ctrl+Space to dictate. Port defaults to 8104; change in the provider's Base URL field if another app is on that port.
Same pattern as Q4 but swaps in cohere-transcribe-q6_k.gguf on port 8105. Q6_K is K-quant mixed-precision, so accuracy sits very close to F16 while the CPU RTFx stays effectively identical to Q4_K (~1.05β1.08Γ on 8 threads per the upstream benchmarks). Use this as the accuracy-first local Cohere; use Q4 only when disk footprint or memory pressure actually matters. The dispatch passes the same backendHint: "cohere" as the Q4 path.
.\scripts\download-cohere-q6k.ps1Or inline:
$dir = "$env:APPDATA\.WhisperInk\cohere-gguf"
New-Item -ItemType Directory -Force -Path $dir | Out-Null
curl.exe -L --fail --progress-bar `
-o "$dir\cohere-transcribe-q6_k.gguf" `
"https://huggingface.co/cstr/cohere-transcribe-03-2026-GGUF/resolve/main/cohere-transcribe-q6_k.gguf"Filename is fixed at cohere-transcribe-q6_k.gguf β the dispatch looks for that exact name, not a glob.
In WhisperInk β Providersβ¦ β pick Cohere Local Q6_K (CrispASR) β Active, Batch. Port defaults to 8105.
Copy the model in, then pick it from the menu. No config file, no rebuild, nothing to start by hand.
1. Put the GGUF in the model folder, %APPDATA%\.WhisperInk\cohere-gguf\ (right-click β π Provider β π Open model folder). The helper script is the safe way to download one: the file only lands in the folder once its size and SHA-256 match what Hugging Face publishes, and an interrupted download resumes.
scripts\get-model.ps1 cstr/orukeet-GGUF # list the repo's .gguf files and sizes
scripts\get-model.ps1 cstr/orukeet-GGUF orukeet-q4_k.gguf # download one2. Right-click the bar or the tray icon β π Provider. Models no provider uses yet are listed under New in the model folder as β name (size). Click one and WhisperInk adds a local provider for exactly that file, gives it a port of its own (8200 and up), switches to it and loads it straight away. A balloon says when it's ready. If your CrispASR can't run it, the balloon says so and WhisperInk switches back to the provider you had.
WhisperInk reads what it needs from the file's own header:
- What it is. CrispASR picks its engine (
--backend) from the file, so there's nothing to set. Files that aren't speech-to-text (voices, punctuation models) aren't offered. - Whether it punctuates. A model with no capitalised words in its vocabulary (Parakeet RNNT 1.1b) writes lowercase with no punctuation, so its provider gets
LocalPuncModel: "fullstop"automatically.
A file that is still being copied shows as β³ until the copy finishes. To remove a model, delete its provider in β Configure Providers (and the file, if you want the space back).
Can your CrispASR run it? crispasr.exe --list-backends lists the engines your binary has (left column), and CrispASR's README lists which models each engine runs. A model newer than your CrispASR needs a CrispASR update first (see CLAUDE.md 5.2, and A/B the new release before you rely on it).
Some models, all from the CrispASR author's Hugging Face repos:
| Model | HuggingFace repo | Good for |
|---|---|---|
parakeet-tdt-0.6b-v3-q4_k.gguf |
cstr/parakeet-tdt-0.6b-v3-GGUF |
Multilingual (25 EU), fast, word timestamps; shipped preset |
orukeet-q4_k.gguf |
cstr/orukeet-GGUF |
A Parakeet TDT 0.6b v3 fine-tune, 402 MB; same engine (CC-BY-SA-4.0) |
parakeet-rnnt-1.1b-q4_k.gguf |
cstr/parakeet-rnnt-1.1b-GGUF |
Stronger English than TDT; lowercase, so punctuation is restored |
canary-1b-v2-q5_0.gguf |
cstr/canary-1b-v2-GGUF |
Explicit-language control + speech translation |
qwen3-asr-1.7b-q4_k.gguf |
cstr/qwen3-asr-1.7b-GGUF |
30 languages + Chinese dialects; shipped preset (port 8112), biasing that works |
granite-speech-4.1-2b-nar-q4_k.gguf |
cstr/granite-speech-4.1-2b-nar-GGUF |
IBM Granite 4.1, non-autoregressive (one pass), 3.4 GB |
voxtral-mini-3b-2507-q4_k.gguf |
cstr/voxtral-mini-3b-2507-GGUF |
Speech-LLM, audio Q&A |
cohere-transcribe-q6_k.gguf |
cstr/cohere-transcribe-03-2026-GGUF |
Strong English, near-F16 at Q6_K; shipped preset |
The menu covers the usual case. To set something it doesn't (a pinned --backend, CPU only, a glob that follows new quants), add an entry to %APPDATA%\.WhisperInk\config.json with WhisperInk closed (it rewrites the file when you change settings), then restart. It shows up under π Provider.
{
"Id": "canary-local",
"Name": "Canary Local (CrispASR)",
"BaseUrl": "http://localhost:8113",
"TranscriptionEndpoint": "http://localhost:8113/v1/audio/transcriptions",
"TranscriberKind": "LocalCrispAsrServer",
"LocalServerPort": 8113,
"LocalModelGlob": "canary-*.gguf",
"LocalBackendHint": "canary",
"BiasMechanism": "none",
"Language": "en"
}Pick a free port. Each local preset spawns its own server, so ports can't be shared. Taken by shipped presets: 8103, 8105β8109, 8112, 8766. Retired but still claimed by old configs: 8102, 8104, 8110, 8111, 8767, 8768. Next free: 8113. Models added from the menu take 8200 and up.
Make LocalModelGlob specific. All presets share one folder and the first filename match wins, so a loose glob will quietly load a different model β which looks like bad accuracy, not a config error. parakeet-*.gguf matches both Parakeet models, hence the pinned parakeet-tdt-* and parakeet-rnnt-1.1b-*. The menu always uses the exact file name.
Optional fields worth knowing:
LocalBackendHintβ not needed: CrispASR v0.8.30 picks the right engine for every model tested, Cohere, Granite and Voxtral included. Set it only to pin a choice.LocalPuncModel: "fullstop"β restores punctuation on models that emit none (Parakeet RNNT/CTC). Skip it for speech-LLM models like Qwen3-ASR, which punctuate themselves.LocalGpuBackendβ blank inherits the global setting; set"cpu"to pin one preset to CPU.BiasMechanism: "hotwords"β a label: the Context Bias terms go to every local model as hotwords either way. They genuinely help Qwen3-ASR and Voxtral 3B (the terms go into the model's prompt), help Granite but can rewrite a correct word, are weak on Parakeet, and are ignored by Cohere and Voxtral 4B.
Test it before relying on it β run the same command WhisperInk will, so a problem is clearly the model's and not your config:
& "$env:APPDATA\.WhisperInk\cohere-gguf\crispasr.exe" --server --host 127.0.0.1 --port 8113 `
-m "$env:APPDATA\.WhisperInk\cohere-gguf\canary-1b-v2-q5_0.gguf" -t 8 -np --backend canary
# then, in another shell:
curl.exe -s http://127.0.0.1:8113/health
curl.exe -s -F "file=@jfk.wav" http://127.0.0.1:8113/v1/audio/transcriptionsA cold CUDA load can take up to two minutes to answer /health β the model is uploaded to VRAM and warmed at startup. That's normal, and only happens once per server.
Cloud providers work the same way if the API speaks the OpenAI multipart shape: same JSON entry, but leave TranscriberKind as "Http" and set BaseUrl, AuthHeaderName (blank means Bearer), ModelFieldName (model or model_id) and TranscriptionModel. APIs with their own protocol β Deepgram, Modulate, Smallest.ai, Reson8, Soniox, Google Chirp 3 β need a small code change instead; see CLAUDE.md β Adding a provider.
Ships as the qwen3-asr-1.7b-local preset β auto-spawned by WhisperInk like the Parakeet and Cohere presets, so there is nothing to start by hand. Drop the GGUF in %APPDATA%\.WhisperInk\cohere-gguf\ and pick the provider:
curl.exe -L -o "$env:APPDATA\.WhisperInk\cohere-gguf\qwen3-asr-1.7b-q4_k.gguf" `
https://huggingface.co/cstr/qwen3-asr-1.7b-GGUF/resolve/main/qwen3-asr-1.7b-q4_k.gguf~1.4 GB. WhisperInk then spawns crispasr.exe --server --backend qwen3-1.7b -m <that file> --port 8112 on first use and keeps it resident.
Two things make this preset different from the other local ones:
- Context-bias terms actually work. Qwen3-ASR is a speech-LLM, so the bias list is spliced into its decoder prompt rather than into a CTC trie. On a clinical test set it recovered
hematochezia(twice) andureterolithiasiswhere the unbiased run produced "hematuria", "hematemesis" and "bursitis with edema" β with no damage to the control clips. Keep the list tight: a 40-term list still fixed 5 of 6 but softenedureterolithiasisinto "ureteral lithiasis". - Punctuation and casing are native, so unlike
parakeet-rnnt-localit needs noLocalPuncModel.
This replaced the old qwen3-asr preset (a plain HTTP entry pointing at a server on localhost:8102 that you had to run yourself). If you still want that arrangement, add a provider with TranscriberKind: "Http" and any inference server exposing /v1/audio/transcriptions.
| File | Purpose |
|---|---|
MainWindow.xaml.cs |
Global low-level keyboard hook, recording state machine, transcription logic, tray & context menu UI, clipboard/paste plumbing. ~1800 LOC β the heart of the app. |
AppConfig.cs |
ApiProvider model (one per backend) and AppConfig (top-level settings, provider list, bias terms). CreateDefaults() seeds the provider list. |
CrispAsrServerTranscriber.cs |
Generic adapter for crispasr.exe --server β one class for every GGUF backend (Cohere, Parakeet, Voxtral, Granite, β¦) via config-only provider entries. |
DeepgramTranscriber.cs |
Deepgram Listen API (/v1/listen) β raw-body POST, Token auth, query-param options, Nova-3 keyterm biasing. |
ModulateTranscriber.cs |
Modulate Velma 2 batch β multipart upload_file, X-API-Key auth, model selected by endpoint path, custom_terms biasing via the JSON config field. |
SmallestTranscriber.cs |
Smallest.ai Waves batch (/waves/v1/stt/) β raw-body POST, Bearer auth, query-param options, transcript at top-level transcription. No biasing surface; language is a strict enum with per-region entitlements, so it is always sent explicitly. |
Reson8Transcriber.cs |
Reson8 prerecorded (/v1/speech-to-text/prerecorded) β raw-body POST, ApiKey auth, query-param options, transcript at top-level text, RFC 7807 errors. Real phrases biasing. language is the mirror image of Smallest's: auto-detect means omitting the param, and the six codes in WhisperInk's dropdown that Reson8 doesn't support are dropped rather than sent (they would 400 every dictation). |
CrispAsrServerTranscriber.cs |
Generic adapter for the crispasr.exe --server mode β model-agnostic, auto-detects backend from GGUF metadata. Used by the Parakeet provider; the path new models should adopt. |
ProviderSettingsWindow.xaml(.cs) |
GUI for editing providers β URLs, keys, auth header, model field, a read-only biasing-mechanism line, and Parakeet hotword-boost / Scribe v2 extra keyterms. |
ContextBiasWindow.xaml(.cs) |
Global context-bias term list β the single source routed to each provider's native biasing field. |
HistoryService.cs, HistoryWindow.xaml(.cs) |
Local transcript log + viewer. |
ApiProvider captures every knob a transcription backend might need:
BaseUrl+ optionalTranscriptionEndpointoverride (full URL).AuthHeaderNameβ blank meansAuthorization: Bearer <key>; set toxi-api-keyfor ElevenLabs; etc.ModelFieldNameβ"model"or"model_id"(ElevenLabs).TranscriptionModel.SupportsTranscription.TranscriptionTemperatureβ nullable; sent as multipart form field when set.BiasMechanism(baked per provider; never user-set) routes the single shared bias list to that provider's native field:"whisper_prompt"β labeled glossary inprompt(OpenAI, local prompt-aware servers)."mistral_context_bias"β comma-joined string incontext_bias(Mistral Voxtral batch, β€100)."elevenlabs_keyterms"β repeatedkeytermsfields (ElevenLabs Scribe v2), sourced from the shared list."hotwords"β comma-joinedhotwordsfor local CrispASR servers (HotwordsBoosttunes the Parakeet trie β off by default, since boosting can garble neighboring words; no-op on Cohere/Granite/Voxtral-4B)."phrase_sets"/"context_terms"/"deepgram_keyterm"/"modulate_custom_terms"/"reson8_phrases"β Google / Soniox / Deepgram / Modulate / Reson8, handled natively in their transcribers. Reson8 is the one where an over-long list degrades accuracy rather than being ignored β keep it tight."none"β provider has no biasing field (e.g. Cohere Transcribe v2, Smallest.ai Waves).
ScribeKeytermsRawβ optional ElevenLabs-only extra keyterms, merged with the shared list and validated together (β€1000 terms, <50 chars, β€5 words, illegal chars dropped).
Default providers (see AppConfig.cs):
| Id | Name | Transport | Notes |
|---|---|---|---|
mistral |
Mistral | HTTPS | Voxtral batch |
openai |
OpenAI | HTTPS | Whisper-1, whisper_prompt bias |
elevenlabs |
ElevenLabs Scribe | HTTPS | xi-api-key, model_id, keyterms |
cohere-api |
Cohere Transcribe API | HTTPS | Cohere v2, temp 0.1; no native biasing |
local |
Local Server | HTTP | localhost:8100, whisper_prompt |
cohere-gguf |
Cohere Local (CrispASR GGUF) | subprocess | llama.cpp CLI |
cohere-gguf-server |
Cohere Local (CrispASR server) | HTTP | llama.cpp server, CPU |
cohere-gguf-cuda-server |
Cohere Local (CrispASR CUDA) | HTTP | llama.cpp server, CUDA |
cohere-gguf-cuda-server-q8 |
Cohere Local (CrispASR CUDA Q8) | HTTP | llama.cpp server, CUDA Q8, cohere_terms |
qwen3-asr-1.7b-local |
Qwen3-ASR 1.7B Local (CrispASR) | HTTP | localhost:8112, auto-spawned --backend qwen3-1.7b; real prompt-splice biasing, native punctuation |
parakeet-local |
Parakeet Local (CrispASR) | HTTP | localhost:8103, auto-spawned via CrispAsrServerTranscriber |
cohere-local-q4 |
Cohere Local Q4 (CrispASR) | HTTP | localhost:8104 β retired from defaults 2026-06-12; still loads from existing configs |
cohere-local-q6k |
Cohere Local Q6_K (CrispASR) | HTTP | localhost:8105, auto-spawned Q6_K Cohere Transcribe (accuracy-first) |
modulate |
Modulate Velma 2 (Multilingual) | HTTPS | X-API-Key, custom_terms biasing |
modulate-english-fast |
Modulate Velma 2 English Fast | HTTPS | Lowest latency; English-only, no biasing |
modulate-multilingual-fast |
Modulate Velma 2 Multilingual Fast | HTTPS | Any language, no metadata, no biasing |
smallest-pulse-pro |
Smallest.ai Pulse Pro (English) | HTTPS | Bearer, raw-body POST; English-only, no biasing |
smallest-pulse |
Smallest.ai Pulse (Multilingual) | HTTPS | Same endpoint, ?model=pulse; 46 languages, no biasing |
reson8 |
Reson8 | HTTPS | ApiKey auth, raw-body POST; phrases biasing (β€250), custom_model_id for larger vocabularies; ten languages |
This table lags
AppConfig.CreateDefaults()β see the provider table inCLAUDE.mdfor the current list.
- Batch β
Clipboard.SetTextfollowed by a syntheticCtrl+V. A single leading space is prepended so the result doesn't fuse to an adjacent word.
SetWindowsHookEx(WH_KEYBOARD_LL, β¦)installed on the UI thread.- Synthetic key presses (the
Ctrl+Vpaste in particular) are tagged with a sentinel flag (0x5AFE) in the hook's extra-info field so the hook ignores its own injections and avoids re-entry. ReleaseAllModifierKeys()runs after every recording to clear any physical modifier that was still held when the hook fired β prevents "stuck Ctrl" after long sessions.
The Cohere v2 transcription endpoint rejects requests where the file part appears before string fields. WhisperInk therefore always appends string fields (model, language, temperature, and any bias field such as context_bias / keyterms) before the file part in the multipart body.
Start/stop chirps are procedurally generated sine waves in memory β no asset files, nothing to ship.
| Path | Contents |
|---|---|
%APPDATA%\.WhisperInk\config.json |
All providers, active id, mic selection, bias terms. |
%APPDATA%\.WhisperInk\debug.log |
This session's log. First place to check for any failure. |
%APPDATA%\.WhisperInk\debug.previous.log |
The previous session's log (kept across one restart). |
%APPDATA%\.WhisperInk\history.json |
Transcription history (viewable from the tray). |
%APPDATA%\.WhisperInk\unsent\ |
Takes that were never delivered (take-*.wav + a .json with the reason). Kept 14 days / 50 takes. |
~/Documents/MyRecordings/temp_audio.wav |
The most recent Batch-mode recording (overwritten each time). |
Config is loaded on startup and rewritten after any settings change. Safe to back up or sync.
# Framework-dependent (smaller; requires the .NET 8 Desktop Runtime on the target machine)
dotnet publish -c Release -r win-x64 --self-contained false
# Self-contained (bundles the runtime; ~80 MB, no prerequisites)
dotnet publish -c Release -r win-x64 --self-contained trueHelper scripts:
publish.ps1β self-contained build into_publish\.publish-framework-dependent.ps1β smaller framework-dependent build into_publish-fd\.scripts\install.ps1 [-Desktop]β runs the self-contained publish then creates Start Menu (and optionally Desktop) shortcuts. The one-shot path from a fresh clone.scripts\install-shortcuts.ps1 [-Desktop]β create shortcuts for an already-published build.scripts\uninstall.ps1 [-RemoveBinaries]β remove shortcuts and the auto-start registry entry. Leaves%APPDATA%\.WhisperInk\alone unless you also pass-RemoveBinaries, which wipes_publish*too.scripts\generate-icon.ps1β regenerateAssets\icon.icofrom code (only needed if you want to change the glyph).
NuGet dependencies (WhisperInk.csproj):
NAudio 2.2.1β microphone capture.Google.Apis.Auth 1.69.0β Google Chirp 3 OAuth (service-account tokens).
If something isn't working, right-click the tray icon and pick Copy support bundle. That drops a zip onto your desktop and puts it on the clipboard so you can paste it straight into Slack / Discord / an issue. The bundle contains:
- The last 500 lines of
%APPDATA%\.WhisperInk\debug.log, and ofdebug.previous.log(the session before) config.jsonwithApiKeyfields redacted (***redacted***)about.txtwith app version, commit hash, .NET and OS versions, installed providers, and which local model files are present
API keys are redacted; GGUF weights are never included (too large).
For a quick live view of what the active provider needs, the tray also has Diagnose active provider β it prints a file/port check block (crispasr.exe FOUND 1.1 MB, port 8104 reachable: YES) so you can see exactly what's in place before you touch anything.
Tray menu quick reference:
- Show Window β restore the floating bar (left-click or double-click the tray icon does the same thing).
- β» Unsent dictations β takes that failed or never finished, each with Retry with the active provider or Retry on a local model (fallback). A recovered transcript goes to the clipboard, not into a window.
- Open debug log / config folder / model folder β opens the paths in Notepad / Explorer.
- Copy support bundle β described above.
- Diagnose active provider β on-demand health probe with per-file detail.
- Aboutβ¦ β version, commit hash, build date.
- View README β opens this page.
- Quit on close β if checked, the X / Alt+F4 actually exits instead of hiding.
- Launch at Windows start β HKCU Run entry; user-level, no admin needed.
- Quit β explicit exit.
App launches but Ctrl+Space does nothing
The keyboard hook needs the main window alive. Don't close it β minimize it. Also check debug.log for hook-install failures (rare; usually means another process is already holding a global hook).
"The .NET runtime is not installed"
Install the .NET 8 Desktop Runtime (x64), not the base runtime. Alternatively, rebuild with --self-contained true.
Recording starts but nothing pastes
Check debug.log for the HTTP response. 401/403 = bad API key. 404 = wrong endpoint (especially common if you customized the TranscriptionEndpoint field). 422 on ElevenLabs with keyterms = swap the repeated form fields for a JSON fallback (commented in MainWindow.xaml.cs at the keyterms block).
Parakeet/Cohere Q4 is slower than expected (RTFx <2Γ on a laptop) Windows' default Balanced power plan throttles CPU to its base clock even while plugged in β on a Ryzen 5825U that's 2.0 GHz versus the 4.5 GHz boost. Roughly halves ASR throughput. Switch to Ultimate Performance:
powercfg -duplicatescheme e9a42b02-d5df-448d-aa00-03f14749eb61
powercfg /setactive 5898ace7-acb8-479d-b9c1-54af0f151d1bFirst command creates the hidden Ultimate Performance scheme from its well-known GUID, second activates it. No reboot needed. Verify with Get-CimInstance Win32_Processor under load β CurrentClockSpeed should now reach MaxClockSpeed under load, not sit at base.
CrispASR/Parakeet returns empty transcripts (or crispasr.exe --help prints nothing)
crispasr.exe is exiting with STATUS_DLL_NOT_FOUND (exit code -1073741515 / 0xC0000135) before it can print anything. One or more DLLs is missing from %APPDATA%\.WhisperInk\cohere-gguf\. Re-run the deploy step and make sure every *.dll gets copied, not just ggml*. Upstream consolidated the per-backend DLLs into crispasr.dll between v0.7.1 and v0.8.30, so a current build needs only crispasr, whisper, ggml, ggml-base, ggml-cpu, ggml-cuda, plus cublas64_12, cublasLt64_12 and cudart64_12 on the CUDA asset. Older builds shipped 13 separate per-backend DLLs (canary, canary_ctc, cohere, granite_speech, parakeet, qwen3_asr, voxtral, voxtral4b, β¦) and need all of them present. Verify with:
$dir = "$env:APPDATA\.WhisperInk\cohere-gguf"
& "$dir\crispasr.exe" --help 2>&1 | Select-Object -First 5
# Exit code 0 and help text = healthy. Exit code -1073741515 = missing DLL.CrispASR build succeeds but no executable appears
Check C:\path\to\CrispASR\build\bin\Release\ β the binary is crispasr.exe, not whisper-cli.exe (CMake renames via OUTPUT_NAME). On Windows there's no whisper-cli symlink; that's a Unix-only step in the upstream CMakeLists.
Stuck modifier key after a recording
Should auto-clear via ReleaseAllModifierKeys(). If it ever happens, tapping the key once releases it. Capture a debug.log excerpt and file an issue.
Mic not listed NAudio enumerates WASAPI devices on launch. Plug in your mic before starting WhisperInk, or restart the app after plugging in.
Longer-form plans and sprint prompts live under plans/:
plans/packaging-polish-prompt.mdβ the specification that drove the tray icon, install scripts, health probe, support bundle, diagnose flow, and auto-start wiring.
PRs welcome. CLAUDE.md has the quick architecture summary that Claude Code uses when working in this repo β a good orientation for new contributors too.