Skip to content

feat: voice pricing — AWS/Azure feeds, a curated floor, and self-refreshing scraped vendor prices - #6

Merged
dittops merged 10 commits into
mainfrom
feat/curated-voice-pricing
Sep 24, 2026
Merged

dittops merged 10 commits into
mainfrom
feat/curated-voice-pricing

Conversation

@dittops

@dittops dittops commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

What

This PR adds prices for the voice vendors bud-connect already offers. It does not add vendors: offering one Bud cannot dispatch is the mistake the fireworks/together_ai withdrawals were about.

No aggregator carries dedicated speech vendors. I checked models.dev, OpenRouter and Helicone, and all three share LiteLLM's blind spot. So the prices come from three sources the SDK now layers over LiteLLM, weakest first:

merge(LiteLLM, ai-models)  ->  curated floor (bundled YAML)  ->  scraped pricing pages  ->  AWS / Azure price feeds
  • A scraped page beats the committed floor because it is fresher.
  • A vendor's own price API beats any page.
  • Every price carries a billing block: the unit, the meter, confidence (authoritative / curated / derived), and the source URL and date. bud-connect stores that block in model_info.billing.

Pairs with BudEcosystem/bud-connect#37, which consumes it. Merge this first. bud-connect pins 583a24e; if this is squash-merged, bump that pin to the merge commit before #37 merges.

1. Authoritative cloud pricing (AWS + Azure)

These feeds are public, unauthenticated and machine-readable. They don't create entries: they attach a billing block (region, feed publication date, confidence: authoritative) to entries LiteLLM already has, and check LiteLLM's rate against the vendor's.

Catalog key Rate Published as
aws_polly/standard / neural / generative / long-form 4e-06 / 1.6e-05 / 3e-05 / 1e-04 per char $4 / $16 / $30 / $100 per 1M
aws_transcribe/StartTranscriptionJob 1e-04/s $0.006 / min
azure/speech/azure-stt 2.78e-04/s $1.00 / hr
azure/speech/azure-tts 1.5e-05/char $15 / 1M

Most of this code exists to avoid four silent failures:

  • Wrong SKU. Transcribe publishes 167 usage types, and Call Analytics costs 5× the base rate while looking just as plausible. The mapping is therefore a hand-verified allow-list.
  • Wrong tier. Volume tiers arrive unordered, so taking min() would bill everyone the top discount. The code selects beginRange == 0 instead.
  • Mangled Azure filter. A raw + in the query string is read as a space, and the server then returns 200 with an empty Items array. The filter is now percent-encoded, and an empty feed logs a warning.
  • Zero-priced "Free" meters. These are refused.

GCP is out of scope because its Billing Catalog refuses unauthenticated callers.

2. Curated floor: 36 entries in data/curated_voice_pricing.yaml

gladia 2 · revai 3 · speechmatics 7 · cartesia 4 · google_speech 13 · deepgram 2 · elevenlabs 3 · assemblyai 2.

  • Each rate was transcribed from the vendor's page, with the URL and date in source.
  • Rates are normalised to per second (STT) or per character (TTS); the vendor's display unit stays in the note.
  • The floor ships inside the package and loads with no network I/O. It is the value that stands whenever a scrape fails.
  • Cartesia sells credits, so its rates are derived: the page's own conversions multiplied by the Pro-tier credit price, never an invented number.
  • Rev AI's "15 second minimum, rounded up to the nearest second" is carried as min_billable_units / rounding_increment.
  • A curated key for a vendor LiteLLM already covers is allowed only for models LiteLLM lacks (the FEED_GAPS test).

3. Scraped pricing: pages that refresh themselves

bud-connect re-seeds every 24h, so a price change on a vendor's page reaches the catalog within a day with nobody in the loop. With no review step, every failure path has to end in "this vendor contributes nothing and the floor stands", never in a wrong number.

Adapter Models Page detail it has to get right
speechmatics 6 STT + TTS Uses the JSON rows (list price), not the rendered promotional table. A row with its own Crossed-out price is read as list, so Linden 1 is $0.40/hr, not the $0.30 promotion.
cartesia sonic-3.6, sonic-3.5, ink-2, ink-whisper Derived from the credit conversions and fails if any of them is missing.
revai 3 Mixes per-hour and per-minute units, and carries the 15s minimum.
gladia async, real-time Starter plan; ignores "as low as" figures.
google_speech_tts / _stt 7 / 6 STT takes the list column, never committed-use savings, and skips the free tier.
deepgram_tts aura-2, aura Pay As You Go, not Growth.
deepgram_stt nova-3, nova-3-general, whisper-large Reads the pre-recorded table (the API WaaV calls), not the streaming table rendered above it.
elevenlabs eleven_v3, eleven_multilingual_v2, eleven_flash_v2_5, eleven_flash_v2, scribe_v2, scribe_v2_medical Flash is absent from LiteLLM. Realtime and deprecated Turbo models are excluded.
assemblyai universal-3-5-pro, universal-2 Reads the async block, not "Universal-3.5 Pro Realtime".

Validation gates, each with a test:

  • Per-unit bounds: 1e-6..1e-2 per second, 1e-7..1e-3 per character.
  • NaN, booleans and zero rates are rejected.
  • A per-adapter min_models floor turns a partial extraction into a failure.
  • A rate more than 10× away from the committed value is rejected as a unit bug.
  • A duplicate row that agrees collapses; one that disagrees rejects the vendor. This is resolved in validate_batch only, so adapters return every match.

Driver, sources/scraped.py:

  • 4 concurrent fetches with a 10s deadline per vendor. It is a real deadline: httpx's timeout bounds each socket read, not the request.
  • A 20s total budget.
  • A page can refresh an existing model's price but never add a model. A key nothing else lists would publish without a drift check.
  • The scraped billing block is merged into the floor's, so rules no page states (Rev AI's rounding) survive a refresh.

Deliberately not scraped:

  • murf and hume: their pages are JS shells with no numbers in the HTML.
  • speechify: no priced model id; the key had been invented, and was reverted.
  • The regional vendors, who quote on contract.

The reasons are written into the YAML header.

4. Catalog correctness fixes found along the way

  • Not models. Entries with no mode (pricing tiers), bedrock/guardrails, deepgram/streaming/* (pricing SKUs), azure/container and openai/container (Code Interpreter sessions), and Together's size-bucket tiers.
  • Retired by the vendor. ElevenLabs scribe_v1* and AssemblyAI best/nano, recorded with evidence in RETIRED_BY_VENDOR.
  • Vertex. LiteLLM's vertex_ai and vertex_ai-embedding-models providers were unmapped, so 16 Gemini-on-Vertex models (TTS, transcription, Live, embeddings) never reached the catalog. All three Vertex providers now feed vertex_ai-gemini-models, filtered to Gemini.
  • Collisions. When two LiteLLM keys map to one catalog key, the canonical entry still wins, but the other now fills its gaps instead of being discarded. The Vertex image models had lost their modalities and search price that way.
  • Model ids. Every key is what budapp sends the vendor as the model name, so it must be an id the vendor's API accepts. That is why Cartesia moved from sonic/ink to its live ids.

Decisions a reviewer should know

  • List price, never promotional. A promotion is recorded in the note, not stored as the rate.
  • No human review gate. A scraped price that passes the gates replaces the floor.
  • Partial success is failure. An adapter that can't extract its whole table contributes nothing.

Testing

  • 324 tests pass. ruff check, ruff format --check and mypy are clean.
  • Every adapter runs against a trimmed fixture of the real page. The new fixtures deliberately keep the regions the adapter must NOT read: Deepgram's streaming table, AssemblyAI's Realtime block, ElevenLabs' realtime cards.
  • Live run on 2026-09-24: 1,554 models, 42 prices from 10/10 vendors, 0 rejected. Every scraped rate was checked by hand against the figure on the vendor's page.
  • Deployed through bud-connect#37 to budconnect-dev and synced into budapp: identical model sets and prices.

🤖 Generated with Claude Code

root and others added 2 commits September 22, 2026 17:00
Workstream B: price the vendors bud-connect already offers. Deliberately NOT
"add vendors" -- every vendor worth pricing is already in the catalog, and
adding one Bud cannot dispatch is the mistake the `fireworks` and `together_ai`
withdrawals were about.

**Curated source.** Every LLM-catalog aggregator shares LiteLLM's blind spot:
dedicated speech vendors are not OpenAI-compatible chat endpoints, so models.dev
(223 providers, 7992 models), OpenRouter and Helicone list none of them. Checked
directly rather than assumed. bud-connect offers those vendors anyway, and a
vendor offered without a price cannot be billed.

So the rates are transcribed by hand from each vendor's pricing page, bundled in
the wheel, loaded with no I/O. 5 models across 2 vendors, each carrying the URL
and the date it was read:

  gladia/async                   $0.61/hr -> 1.6944e-04/s
  gladia/real-time               $0.75/hr -> 2.0833e-04/s
  revai/reverb                   $0.20/hr -> 5.5556e-05/s
  revai/reverb-foreign-language  $0.30/hr -> 8.3333e-05/s
  revai/whisper-large            $0.005/MIN -> 8.3333e-05/s

Rev AI quotes Reverb per hour and Whisper Large per minute -- two units inside
one vendor -- which is why the file normalises everything to per-second and
records the vendor's display unit in `source.note`. It also bills "rounded up to
the nearest second, 15 second minimum", carried as min_billable_units/
rounding_increment for bud-connect's new `billing` block.

Only verified numbers are here. Cartesia sells credits and Speechmatics
publishes one unattributed figure, so both are left out with the reason written
down; bud-connect stamps confidence=UNKNOWN on anything unpriced, and a consumer
that sees UNKNOWN can refuse to bill, while one that sees an invented number
cannot tell it was invented.

**AWS Transcribe.** LiteLLM prices it under `transcribe`; bud-connect has
offered `aws_transcribe` all along with no models behind it. One mapping line
each in LITELLM_TO_TENSORZERO, TENSORZERO_PROVIDERS and STRIP_PREFIXES -- the
last keyed by the TensorZero name, not the LiteLLM one, or it silently no-ops
and the URI becomes `aws_transcribe/transcribe/StartTranscriptionJob`. Test
covers exactly that.

**The overlay is additive and refuses to run on an empty merge.** A live-feed
key wins over a curated one, with a warning saying the curated entry is now
redundant. And an empty merge means the upstream fetch failed: overlaying there
would turn a total outage into a catalog of five voice models that looks like a
successful sync, and bud-connect's seeder would retire every other model on the
strength of it. `test_end_to_end_empty_litellm` caught that.

20 new tests, 90 total. ruff and mypy clean. Wheel verified to contain the YAML.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Workstream B2. AWS Price List and Azure Retail Prices are public, unauthenticated
and machine-readable, which makes them the only speech prices in this catalog
that refresh themselves. GCP is deliberately absent: its Cloud Billing Catalog
refuses unauthenticated callers and this SDK carries no credentials.

These sources do not invent catalog entries. LiteLLM already lists
`aws_polly/standard`, `azure/speech/azure-stt` and the rest; this attaches the
`billing` block they cannot carry -- region, the vendor's own publication date,
`confidence: authoritative` -- and audits the rate while it is there.

The audit is the part worth having. Against live data today all 7 canonical
rates agree with LiteLLM to the last digit:

  aws_polly/standard                    4e-06/char    ($4/1M)
  aws_polly/neural                      1.6e-05/char  ($16/1M)
  aws_polly/generative                  3e-05/char    ($30/1M)
  aws_polly/long-form                   1e-04/char    ($100/1M)
  aws_transcribe/StartTranscriptionJob  1e-04/s       ($0.006/min)
  azure/speech/azure-stt                2.777778e-04/s ($1.00/hr)
  azure/speech/azure-tts                1.5e-05/char  ($15/1M)

A future disagreement means the vendor repriced, LiteLLM drifted, or the SKU is
mapped to the wrong entry. All three want a human, so the vendor's number wins
but never quietly -- both values go in the warning.

THE SILENT FAILURES THIS IS BUILT AROUND

* Wrong SKU. Transcribe publishes 167 usage types;
  `USE1-CallAnalyticsTranscribeAudio` is 5x `USE1-TranscribeAudio` and reads just
  as plausibly. The mapping is a hand-verified allow-list; anything else is
  ignored. Call Analytics, Redaction, Medical, HealthScribe, Toxicity and the
  custom-language-model variants are all excluded on purpose.
* Wrong tier. AWS lists volume tiers as unordered price dimensions under one
  term. `min()` would bill everyone the highest-volume discount, so selection is
  by `beginRange == 0` -- list price.
* Mangled filter. Azure's service family is literally "AI + Machine Learning",
  and a raw `+` in a query string is a SPACE to the server: the filter becomes
  "AI   Machine Learning", matches nothing, and returns HTTP 200 with an empty
  Items array. Found exactly this way -- the first live run resolved 0 meters and
  raised nothing. Now percent-encoded, and an empty feed warns.
* Zero rates. Azure publishes "Free <meter>" rows beside the paid ones. A zero
  reaching the catalog is a customer billed nothing, silently; those rows are
  refused.

`Neural HD Text to Speech Characters` ($22/1M) is deliberately NOT mapped:
LiteLLM's `azure/speech/azure-tts-hd` is $30/1M and the two are probably
different products. Mapping them would overwrite a possibly-correct rate with a
confidently wrong one.

Both sources are best-effort, like ai-models. They are two more network calls on
a nightly sync whose failure mode is bud-connect retiring models; losing a
billing block is recoverable next run, losing the catalog is not.

11 new tests, 101 total. ruff and mypy clean. All 12 billing blocks the SDK now
emits (7 authoritative, 5 curated) validate against bud-connect's Billing schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@dittops dittops changed the title feat: curated voice pricing, and map AWS Transcribe feat: voice pricing — curated vendors, and authoritative AWS/Azure feeds Sep 22, 2026
root and others added 8 commits September 23, 2026 17:20
The previous commit left both out claiming there was nothing to transcribe. That
was wrong, and the reason is worth recording: I judged them from a
markdown-converted summary of the page rather than the page. The raw HTML has more
in it.

**speechmatics** (7 models, curated). The page embeds its pricing table as
structured JSON (`"Section":"Pricing"` rows). I had reported "one figure, $0.129,
without saying which model" -- in fact there is a full per-model table:

  Batch Melia 1        $0.24/hr -> 6.6667e-05/s
  Batch Standard       $0.45/hr -> 1.25e-04/s
  Batch Enhanced       $0.75/hr -> 2.0833e-04/s
  Real-time Standard   $0.45/hr -> 1.25e-04/s
  Real-time Enhanced   $0.80/hr -> 2.2222e-04/s
  Linden 1             $0.30/hr -> 8.3333e-05/s
  Text-to-Speech       $0.011/1k chars -> 1.1e-05/char

The JSON carries LIST prices while the rendered page shows a promotional rate ~46%
lower (Melia 1 displays $0.129/hr against $0.24 list). List is stored, for the
same reason AWS's tier-0 rate is: a promotion expires, and over-estimating a cost
is the safer error. Each note records the promotional figure so the gap is visible.
Bolt-ons (Translation, Summaries, Chapters, Sentiment, Topics) are per-feature
surcharges, not models, and are excluded like AWS's Call Analytics SKUs.

**cartesia** (2 models, DERIVED). It does sell credits, so "no published per-unit
price" was right -- but it publishes every conversion needed to compute one, which
is the difference between deriving and inventing:

  "1 credit equals 1 character"           -> sonic  6.5e-05/char
  "3 credits equals 1 second of audio"    -> ink    1.95e-04/s
  "At pro tier, it's $65 per 1M credits"  (startup $45, scale $38)

Pro is the entry paid tier, so it is list. `confidence: derived` is the label that
exists for exactly this: the arithmetic is an assumption about how Cartesia
converts credits and breaks if they reprice a tier without repricing credits. The
note carries a cross-check -- Cartesia states Ink is $0.39/hr on Scale, the
conversion gives $0.41/hr, a 5% gap -- which is why it is derived and not curated.

Tests generalised rather than special-cased: `unit` must now agree with whichever
cost field an entry carries, per-character rates get their own round-trip check,
Cartesia's credit arithmetic is recomputed in a test so a typo fails instead of
becoming a bill, and `derived` is asserted to stay confined to credit-based vendors
so it cannot become a soft label for a rate nobody checked.

Still absent, with the reason recorded in the file: unrealspeech (character quotas,
no per-unit rate exists), murf and hume (JS-shell pricing pages, no numbers in the
HTML at all), resemble (the /pricing/ page sells deepfake detection, not TTS).

Curated total 5 -> 14. Catalog billing blocks: 7 authoritative, 12 curated,
2 derived. 111 tests, ruff and mypy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Vendors like Speechmatics and Cartesia publish prices only as a web page. Until now
that meant I read the page and typed the number into a YAML file, and the
`checked_on` date advanced only when someone remembered to look. This adds adapters
that read those pages directly, so bud-connect's existing 24h `cron-tensorzero-sync`
carries a price change into the catalog on its own.

No new scheduler, and nothing in CI: bud-connect already re-seeds every 24 hours via
a dapr cron binding, verified live in the budconnect-dev install. The SDK just
provides the data; the refresh was already solved.

WHAT MAKES THIS SAFE WITHOUT REVIEW

There is no human between a page changing and a price being served, so every failure
path is designed to end in *the vendor contributing nothing*, never in a wrong number
and never in a missing one:

  page unreachable / non-200 / slow   -> vendor skipped
  adapter raises                      -> vendor skipped, siblings unaffected
  fewer models than the adapter's floor -> nothing taken from that vendor
  a single rate fails a gate          -> that model dropped, the rest kept
  source overruns its 20s budget      -> whatever finished is used
  base catalog empty                  -> overlay skipped entirely

In each case the committed rate in curated_voice_pricing.yaml stands, so scraping can
only ever make the catalog staler -- not wrong, not emptier. That matters because
bud-connect's seeder reads "absent from this run" as "deactivate", which makes a
partial scrape more damaging than no scrape.

THE GATES, which are the only guard left once review is gone

Bounds per unit (per-second 1e-6..1e-2, per-character 1e-7..1e-3) catch the failures
scraping actually produces: a per-hour figure stored as per-second is 3600x out, a
misplaced decimal is 10x or 100x. Non-finite rates are rejected explicitly because
NaN compares False against every bound and would otherwise pass every range check.
`True` is rejected as a rate because it is an int in Python and would pass `> 0`.
Duplicate keys reject the whole extraction, since which rate wins would be an
accident of ordering and one duplicated row makes the rest of the parse suspect.

Above that, the overlay rejects a rate that has moved more than 10x from the
committed floor. Vendors halve and double prices -- Speechmatics currently shows a
46% promotional cut -- but they do not move by factors of ten, while a parser reading
per-hour as per-second does. An authoritative cloud price is never overwritten by a
marketing page, and an entry a live feed priced on a different basis is left alone
rather than left with two cost fields disagreeing.

PRECEDENCE

  merge(litellm, ai_models) -> curated floor -> scraped -> authoritative cloud

Scraped beats the committed floor, so a new price is taken with no review. A vendor's
own price API still beats any page.

ADAPTERS

speechmatics (7 models) reads the pricing table the page embeds as JSON, taking LIST
prices and recording the displayed promotion in the note rather than using it.
Bolt-ons (Translation, Summaries, Chapters, Sentiment, Topics) are excluded as
per-feature surcharges, like AWS's Call Analytics SKUs.

cartesia (2 models, `derived`) parses all three stated conversions -- "1 credit
equals 1 character", "3 credits equals 1 second of audio", "at pro tier, it's $65 per
1M credits" -- and fails if any is missing. Nothing is hardcoded, because a hardcoded
$65/1M would keep producing plausible numbers long after Cartesia changed it, which
is exactly how a derived rate goes quietly wrong.

Both were checked against the live pages: 9 prices, 0 rejected, every rate identical
to the value I had transcribed by hand.

TESTS: 99 new, 210 total

Fixtures are trimmed regions of the real pages (27KB and 1.4KB, from 577KB and
192KB) and reproduce the full-page extraction exactly, promotional figures included.
No test touches the network -- an adapter that needed a live page to be tested would
fail during a vendor's outage. One test asserts each adapter still agrees with the
committed floor, because that floor is the fallback used when a scrape fails, and a
silent divergence would make the fallback wrong in precisely the situation it exists
for.

Also: pages are fetched with an honest `bud-model-catalog/1.0` user agent, verified to
return identical content to a browser UA, so there is no reason to impersonate one.
Concurrency is bounded so eight vendors are not hit simultaneously every 24 hours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…arges

Adds the two remaining adapters for vendors already in the curated file, so all 14
hand-curated prices are now self-refreshing rather than dated by whenever someone
last looked. `python -m bud_model_catalog.scrapers --diff` reports 14 prices from 4
vendors with no divergence from the committed file at all.

A DATA-LOSS BUG THIS FOUND

Rev AI publishes "Rounded up to the nearest second, 15 second minimum" on every
model, and the curated entry carried `min_billable_units: 15`. The overlay replaces
an entry's whole billing block, so scraping Rev AI would have written a block without
that field and silently dropped it -- leaving a rate that under-bills every clip
shorter than 15 seconds, with nothing in a diff to show it had happened. Short clips
are the common case for voice, so this would have been a real and invisible
under-billing.

`min_billable_units` now travels the whole path: extracted by the adapter, validated
(positive, finite, under an hour -- a "minimum" of thousands of seconds is a parsed
year, not a billing floor), carried into the entry, and asserted by a regression test.
Absent means the key is omitted rather than set to null, because null reads as
"charges from zero", which is a different claim from "the vendor publishes no
minimum".

REV AI IS PARSED BY STRUCTURE, NOT BY BLOB

The first version regexed a tag-stripped page and produced a model named
"supports-all-popular-media-types-email-and-chat-support-get-started-view-more-offerings-reverb",
because stripping tags joins every element into one line and a non-greedy match had
no boundary to stop at. It now walks the page's own `payment-offering` blocks, so a
name cannot absorb the sentence in front of it. There is a test pinning that.

Excluded: Human Transcription ($1.99/min, which is people), and Forced Alignment,
Language Identification, Translation, Sentiment, Summarization and Topic Extraction,
which are applied to a transcript rather than deployed -- the same category as AWS's
Call Analytics SKUs and Speechmatics' bolt-ons. Rev AI's "per 10 words" units are
skipped rather than mis-stored, since the catalog cannot express them.

GLADIA TAKES THE LIST PRICE

Starter is "Async at $0.61 /hr"; Growth is "as low as $0.20 /hr" and needs an upfront
commitment, making it a volume discount rather than a list price. The pattern anchors
on "at $" so it cannot pick up a committed rate by accident, and the fixture keeps the
Growth card specifically so a test can prove it does not.

226 tests, ruff and mypy clean. Fixtures total 43KB against 1.1MB of original pages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…wo cannot be

Of the four remaining vendors I said were "static and scrapable", two hold up and two
do not. That earlier classification counted dollar amounts in the HTML without
checking whether they were per-unit rates for the vendor's own models, and it was too
optimistic.

ADDED

google_speech (7 models) is the valuable one: four capabilities and zero models in the
catalog until now, and the only vendor in the set that quotes the catalog's own unit
directly -- "US$0.00003 per character", no conversion to get wrong. Chirp 3 HD $30/1M,
instant custom voice $60/1M, WaveNet and Standard $4/1M, Studio $160/1M, Neural2 and
Polyglot $16/1M. Parsed by table row rather than from stripped text, for the reason
the Rev AI adapter records.

  Gemini TTS rows are skipped: they bill per million input text tokens and output audio
  tokens, which is a different basis, and storing one as a character rate would be
  wrong by orders of magnitude in a way the bounds check would NOT catch. The fixture
  keeps a Gemini row so a test holds that line.

  Google's speech-to-TEXT page is deliberately not covered. Its table interleaves list
  prices with 1-year and 3-year committed-savings columns, so reading a rate means
  knowing which column you are in, and getting it wrong under-bills by 20% with nothing
  to show it. That needs an adapter written against the structure, not a guess.

speechify (1 entry) charges per character across all voices at a per-plan rate, naming
no priced model, so this is a single `text-to-speech` entry -- a stand-in key following
the label Speechmatics' own table uses, NOT a deployable model name. Worth knowing
before it is shown to anyone as a model. $10/1M entry plan is stored as list; $8 Pro
and $6 Scale are volume tiers.

  Its page carries the nastiest trap in this whole exercise: a third-party comparison
  of COMPETITORS' prices per million characters, including Cartesia at $49/1M. A
  regex for "$X per million characters" finds all of them. Reading Cartesia's rate off
  Speechify's marketing page would be absurd on its own, and a future adapter matching
  on vendor names could have overwritten the rate this SDK derives from Cartesia's own
  page. The pattern anchors on the plan-overage phrasing instead, and a test feeds the
  comparison table in and asserts nothing comes out.

NOT ADDED, with the evidence recorded in the YAML header

elevenlabs: 710KB of HTML holding the pricing table's LABELS
("price_per_1k_characters": "Price per 1K characters") and none of its values, which
render client-side. What it does state is credits, marked "approximate", with per-model
variation given as a range ("between 0.5 and 1 credit per character") and no
credit-to-dollar conversion anywhere. Even with a browser that is a range, not a rate.

smallest: prices voice-agent layers approximately ("Text to Speech Layer
~$0.09/minutes"). The only per-model figures on the page belong to third-party LLMs, so
extracting them would attribute OpenAI's prices to smallest. Its own models are never
priced by name.

TWO BUGS FOUND WHILE DOING IT

YAML reads `3e-05` as a string, not a float -- scientific notation needs a decimal
point and signed exponent (`3.0e-05`) to parse as a number. Every generated entry for
the two new vendors arrived as a string, and the failure surfaced as a TypeError
comparing str to int in an unrelated test. Rates are now written in plain decimal, and
CuratedSource rejects a non-numeric rate with a message that says what to do, so the
next person to hand-write a small number gets told rather than debugging a stack trace.

Two scrapers registered for one vendor could produce the same catalog key and silently
overwrite each other -- `validate_batch` catches duplicates within one extraction but
not across two. Exactly how Google would look with both its pages registered and both
yielding a "standard". The clash is now logged as the registry bug it is and the first
value kept.

22 prices from 6 vendors, no divergence from the committed floor. 250 tests, ruff and
mypy clean. Catalog billing blocks: 7 authoritative, 20 curated, 2 derived.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…0% error

Six models: recognition $0.016/min, dynamic batch $0.003, speech recognition with
data logging $0.016 and without $0.024, medical dictation and conversation $0.078.
google_speech now carries 13 models -- 7 text-to-speech, 6 speech-to-text -- where it
had none two commits ago.

WHY THIS ONE WAITED

Every row on this page has three price columns, and in stripped text they are
indistinguishable:

  Recognition (sku:3099-B70F-0949) | Standard
    | 0 to 500,000 min  $0.016 / 1 minute ...    <- list
    | 0 to 500,000 min  $0.0144 / 1 minute ...   <- 1-year commitment
    | 0 to 500,000 min  $0.0128 / 1 minute ...   <- 3-year commitment

Storing $0.0144 instead of $0.016 under-bills by 10% and is undetectable: the number
is plausible, the unit is right, the magnitude is right, and neither the bounds check
nor the drift check would ever flag it. Pattern-matching the page the way the other
six adapters do would have been a coin flip, which is why this was left out of the
previous commit rather than guessed at.

Three defences instead:

  * The list column is located by its HEADING, never by position, and a table whose
    heading cannot be identified yields nothing. There is deliberately no positional
    fallback -- "the third column is usually list" is the entire risk.
  * Having taken a rate, the adapter asserts it is not cheaper than the same row's
    remaining columns. List is by definition dearer than a rate you commit a year to,
    so if it is not, the columns were misread and the table is rejected outright.
  * The free tier each cell opens with ("0 minute to 60 minute $0.00 (Free)") is
    skipped and the first paid tier taken -- the lowest-volume, undiscounted rate, the
    same rule applied to AWS tier 0. Taking the first number in the cell instead would
    price these models at zero, which is a request billed as free and looks
    unremarkable in a catalog.

Unlike the TTS names, parentheses are preserved: "with data logging" and "without data
logging" differ by 50% in price, and the distinction lives entirely inside the
parenthetical.

FRAMEWORK: MORE THAN ONE ADAPTER PER VENDOR

The registry was a dict keyed by vendor, which cannot express "Google publishes TTS and
STT on separate pages with different table shapes". It is now a tuple, and adapters
carry an optional `name` distinct from `vendor` (`google_speech_tts`,
`google_speech_stt`) used to select one on the command line and to find its fixture.
`vendor` remains the catalog key. The cross-adapter key-collision guard added last
commit now has a real case to guard, and a test asserts the two Google adapters claim
no key in common.

The fixture keeps all four pricing tables WITH their savings columns and free tiers,
because those are precisely what the adapter has to avoid reading -- a fixture without
them would test nothing. One test enumerates every committed-savings rate on the page
and asserts none of them was ever stored; another feeds a table whose columns are
swapped and asserts the invariant fires.

28 prices from 7 adapters, no divergence from the committed floor. 268 tests, ruff and
mypy clean. Catalog billing blocks: 7 authoritative, 26 curated, 2 derived.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Speechify meters per character across all voices at a per-plan rate and names no
priced model anywhere on its page. The rate is real; there is nothing to attach it to.
`speechify/text-to-speech` was a label this project made up, and it would have appeared
in budapp's model list as something nobody can deploy -- the same objection that had
fireworks and together_ai withdrawn from the provider list.

Removes the adapter, its fixture, its tests and its floor entry, and records the reason
in the YAML header alongside the other absences. That note also carries the warning for
whoever revisits it: Speechify's page lists third-party estimates of COMPETITORS'
rates, Cartesia at $49/1M among them, so a future adapter must not match on
"$X per million characters".

27 curated prices across 5 vendors, 6 adapters. 261 tests, ruff and mypy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… models

DEEPGRAM ADVERTISED TEXT-TO-SPEECH WITH NO TTS MODEL

LiteLLM carries 42 Deepgram entries and every one is audio_transcription, so Aura never
reached the catalog. `scrapers/vendors/deepgram.py` supplies it from deepgram.com/pricing:

  aura-2   $0.030 / 1k characters   (90 voices)
  aura     $0.015 / 1k characters   (12 voices; the pricing page calls it "Aura-1")

Priced per family, not per voice, because that is how Deepgram bills -- the same
granularity as Google's WaveNet and Studio. Keys are the architecture names Deepgram's own
/v1/models uses, not the marketing names, so `aura` rather than `aura-1`. Pay As You Go is
list; Growth is recorded in the note.

Flux TTS is priced on the same page and deliberately NOT emitted. It is absent from
/v1/models, and the only `flux-tts` string in Deepgram's docs is a sidebar URL slug whose
pages return 404, so nothing establishes what `model=` value selects it. A key invented here
would reach budapp as a model nobody can deploy, which is why Speechify was removed.

The curated floor now carries Deepgram. The test that forbade curating a feed-covered
vendor was vendor-level, correct while coverage was all-or-nothing; it is now
(vendor, mode) -- a fed vendor may be curated only in a mode the feed does not supply --
which keeps the intent while letting a genuine gap be filled. If LiteLLM adds Aura later,
the runtime overlay lets the feed win and logs that the curated entry should go.

FIFTEEN ENTRIES THAT WERE NEVER MODELS

`not_a_model()` excludes them at the source. Each rule was checked against the live price
map and removed exactly these and nothing else:

  8  no mode       fireworks-ai-4.1b-to-16b and siblings: size-bucket pricing tiers.
                   With no mode there is nothing to derive a route from.
  1  guardrail     bedrock/guardrails: a filtering product applied to other models.
  6  deepgram      deepgram/streaming/*: Deepgram's /v1/models has no streaming model and
                   no `nova-3-multilingual` -- streaming is a way of calling the nova-3
                   models already listed, and multilingual is nova-3 with language=multi.
                   detect_entities, diarize, keyterm and redact are feature surcharges.

All fifteen reached bud-connect with no endpoint, and budapp offered them as deployable.
Realtime models are deliberately NOT excluded: they are real models, and whether Bud can
route them is the consumer's question, not a reason to erase them.

29 prices from 7 adapters, 0 rejected. 285 tests; ruff and mypy clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…osing fields

A review of every provider's catalog data found entries that were wrong in ways no test
caught. Each is fixed here with a test that fails without the fix.

Prices
- Speechmatics Linden 1 was stored at its $0.30/hr promotion. Its JSON row carries its own
  "Crossed-out price" of 0.40 ("discounted 25% from a $0.40/hr list rate"); the adapter now
  reads that as list, and the floor is corrected.
- The scraped overlay replaced the whole billing block, deleting floor-only rules such as
  Rev AI's rounding_increment from all three Rev AI models. It now merges.
- A scraped key nothing else lists was published unchecked: a renamed row at 10x became a
  second model beside the stale first. Pages now refresh prices only; new models enter
  through the curated file. Slugs keep '.', so "Linden 1.1" is not "linden-11".
- Every adapter kept "the first" of a repeated row itself, which made validate_batch's
  duplicate rejection unreachable. Duplicates are now resolved there only: a repeat that
  agrees collapses, one that disagrees rejects the vendor.
- httpx's timeout bounds a read, not a request, so a dribbling server held its slot for the
  whole budget. Each vendor now has a real deadline.

Model ids (the key is what budapp sends the vendor as the model)
- Cartesia `sonic` was sunset on 2026-06-01 and `ink` was never an id. Keyed now by the live
  ids: sonic-3.6, sonic-3.5, ink-2, ink-whisper.
- ElevenLabs scribe_v1/scribe_v1_experimental and AssemblyAI best/nano are withdrawn from the
  vendors' APIs; dropped, with the evidence recorded.
- azure/container, openai/container (Code Interpreter sessions) and Together's size-bucket
  tiers were published as models; excluded.

Coverage
- LiteLLM's `vertex_ai` and `vertex_ai-embedding-models` providers were unmapped, so 16 Gemini
  models on Vertex (TTS, transcription, Live, embeddings, gemini-omni) never reached the
  catalog. All three Vertex providers now feed vertex_ai-gemini-models, Gemini only.
- Colliding LiteLLM keys kept one entry wholesale and discarded the other's fields -- the
  Vertex image models lost their modalities and search price. The canonical entry still
  wins; the other now fills its gaps.
- New scrapers: ElevenLabs (Flash, which LiteLLM lacks, plus Multilingual, v3, Scribe v2 and
  v2 Medical), Deepgram pre-recorded STT (Nova-3; Whisper Large, which LiteLLM had 25% high),
  and AssemblyAI (universal-3-5-pro, universal-2, replacing the retired ids).

Live run: 1550 models, 42 prices scraped from 10/10 vendors, 0 rejected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dittops
dittops marked this pull request as ready for review September 24, 2026 13:38
@dittops dittops changed the title feat: voice pricing — curated vendors, and authoritative AWS/Azure feeds feat: voice pricing — AWS/Azure feeds, a curated floor, and self-refreshing scraped vendor prices Sep 24, 2026
@dittops
dittops merged commit 943f4c0 into main Sep 24, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant