feat: voice pricing — AWS/Azure feeds, a curated floor, and self-refreshing scraped vendor prices - #6
Merged
Merged
Conversation
Workstream B: price the vendors bud-connect already offers. Deliberately NOT "add vendors" -- every vendor worth pricing is already in the catalog, and adding one Bud cannot dispatch is the mistake the `fireworks` and `together_ai` withdrawals were about. **Curated source.** Every LLM-catalog aggregator shares LiteLLM's blind spot: dedicated speech vendors are not OpenAI-compatible chat endpoints, so models.dev (223 providers, 7992 models), OpenRouter and Helicone list none of them. Checked directly rather than assumed. bud-connect offers those vendors anyway, and a vendor offered without a price cannot be billed. So the rates are transcribed by hand from each vendor's pricing page, bundled in the wheel, loaded with no I/O. 5 models across 2 vendors, each carrying the URL and the date it was read: gladia/async $0.61/hr -> 1.6944e-04/s gladia/real-time $0.75/hr -> 2.0833e-04/s revai/reverb $0.20/hr -> 5.5556e-05/s revai/reverb-foreign-language $0.30/hr -> 8.3333e-05/s revai/whisper-large $0.005/MIN -> 8.3333e-05/s Rev AI quotes Reverb per hour and Whisper Large per minute -- two units inside one vendor -- which is why the file normalises everything to per-second and records the vendor's display unit in `source.note`. It also bills "rounded up to the nearest second, 15 second minimum", carried as min_billable_units/ rounding_increment for bud-connect's new `billing` block. Only verified numbers are here. Cartesia sells credits and Speechmatics publishes one unattributed figure, so both are left out with the reason written down; bud-connect stamps confidence=UNKNOWN on anything unpriced, and a consumer that sees UNKNOWN can refuse to bill, while one that sees an invented number cannot tell it was invented. **AWS Transcribe.** LiteLLM prices it under `transcribe`; bud-connect has offered `aws_transcribe` all along with no models behind it. One mapping line each in LITELLM_TO_TENSORZERO, TENSORZERO_PROVIDERS and STRIP_PREFIXES -- the last keyed by the TensorZero name, not the LiteLLM one, or it silently no-ops and the URI becomes `aws_transcribe/transcribe/StartTranscriptionJob`. Test covers exactly that. **The overlay is additive and refuses to run on an empty merge.** A live-feed key wins over a curated one, with a warning saying the curated entry is now redundant. And an empty merge means the upstream fetch failed: overlaying there would turn a total outage into a catalog of five voice models that looks like a successful sync, and bud-connect's seeder would retire every other model on the strength of it. `test_end_to_end_empty_litellm` caught that. 20 new tests, 90 total. ruff and mypy clean. Wheel verified to contain the YAML. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Workstream B2. AWS Price List and Azure Retail Prices are public, unauthenticated and machine-readable, which makes them the only speech prices in this catalog that refresh themselves. GCP is deliberately absent: its Cloud Billing Catalog refuses unauthenticated callers and this SDK carries no credentials. These sources do not invent catalog entries. LiteLLM already lists `aws_polly/standard`, `azure/speech/azure-stt` and the rest; this attaches the `billing` block they cannot carry -- region, the vendor's own publication date, `confidence: authoritative` -- and audits the rate while it is there. The audit is the part worth having. Against live data today all 7 canonical rates agree with LiteLLM to the last digit: aws_polly/standard 4e-06/char ($4/1M) aws_polly/neural 1.6e-05/char ($16/1M) aws_polly/generative 3e-05/char ($30/1M) aws_polly/long-form 1e-04/char ($100/1M) aws_transcribe/StartTranscriptionJob 1e-04/s ($0.006/min) azure/speech/azure-stt 2.777778e-04/s ($1.00/hr) azure/speech/azure-tts 1.5e-05/char ($15/1M) A future disagreement means the vendor repriced, LiteLLM drifted, or the SKU is mapped to the wrong entry. All three want a human, so the vendor's number wins but never quietly -- both values go in the warning. THE SILENT FAILURES THIS IS BUILT AROUND * Wrong SKU. Transcribe publishes 167 usage types; `USE1-CallAnalyticsTranscribeAudio` is 5x `USE1-TranscribeAudio` and reads just as plausibly. The mapping is a hand-verified allow-list; anything else is ignored. Call Analytics, Redaction, Medical, HealthScribe, Toxicity and the custom-language-model variants are all excluded on purpose. * Wrong tier. AWS lists volume tiers as unordered price dimensions under one term. `min()` would bill everyone the highest-volume discount, so selection is by `beginRange == 0` -- list price. * Mangled filter. Azure's service family is literally "AI + Machine Learning", and a raw `+` in a query string is a SPACE to the server: the filter becomes "AI Machine Learning", matches nothing, and returns HTTP 200 with an empty Items array. Found exactly this way -- the first live run resolved 0 meters and raised nothing. Now percent-encoded, and an empty feed warns. * Zero rates. Azure publishes "Free <meter>" rows beside the paid ones. A zero reaching the catalog is a customer billed nothing, silently; those rows are refused. `Neural HD Text to Speech Characters` ($22/1M) is deliberately NOT mapped: LiteLLM's `azure/speech/azure-tts-hd` is $30/1M and the two are probably different products. Mapping them would overwrite a possibly-correct rate with a confidently wrong one. Both sources are best-effort, like ai-models. They are two more network calls on a nightly sync whose failure mode is bud-connect retiring models; losing a billing block is recoverable next run, losing the catalog is not. 11 new tests, 101 total. ruff and mypy clean. All 12 billing blocks the SDK now emits (7 authoritative, 5 curated) validate against bud-connect's Billing schema. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The previous commit left both out claiming there was nothing to transcribe. That was wrong, and the reason is worth recording: I judged them from a markdown-converted summary of the page rather than the page. The raw HTML has more in it. **speechmatics** (7 models, curated). The page embeds its pricing table as structured JSON (`"Section":"Pricing"` rows). I had reported "one figure, $0.129, without saying which model" -- in fact there is a full per-model table: Batch Melia 1 $0.24/hr -> 6.6667e-05/s Batch Standard $0.45/hr -> 1.25e-04/s Batch Enhanced $0.75/hr -> 2.0833e-04/s Real-time Standard $0.45/hr -> 1.25e-04/s Real-time Enhanced $0.80/hr -> 2.2222e-04/s Linden 1 $0.30/hr -> 8.3333e-05/s Text-to-Speech $0.011/1k chars -> 1.1e-05/char The JSON carries LIST prices while the rendered page shows a promotional rate ~46% lower (Melia 1 displays $0.129/hr against $0.24 list). List is stored, for the same reason AWS's tier-0 rate is: a promotion expires, and over-estimating a cost is the safer error. Each note records the promotional figure so the gap is visible. Bolt-ons (Translation, Summaries, Chapters, Sentiment, Topics) are per-feature surcharges, not models, and are excluded like AWS's Call Analytics SKUs. **cartesia** (2 models, DERIVED). It does sell credits, so "no published per-unit price" was right -- but it publishes every conversion needed to compute one, which is the difference between deriving and inventing: "1 credit equals 1 character" -> sonic 6.5e-05/char "3 credits equals 1 second of audio" -> ink 1.95e-04/s "At pro tier, it's $65 per 1M credits" (startup $45, scale $38) Pro is the entry paid tier, so it is list. `confidence: derived` is the label that exists for exactly this: the arithmetic is an assumption about how Cartesia converts credits and breaks if they reprice a tier without repricing credits. The note carries a cross-check -- Cartesia states Ink is $0.39/hr on Scale, the conversion gives $0.41/hr, a 5% gap -- which is why it is derived and not curated. Tests generalised rather than special-cased: `unit` must now agree with whichever cost field an entry carries, per-character rates get their own round-trip check, Cartesia's credit arithmetic is recomputed in a test so a typo fails instead of becoming a bill, and `derived` is asserted to stay confined to credit-based vendors so it cannot become a soft label for a rate nobody checked. Still absent, with the reason recorded in the file: unrealspeech (character quotas, no per-unit rate exists), murf and hume (JS-shell pricing pages, no numbers in the HTML at all), resemble (the /pricing/ page sells deepfake detection, not TTS). Curated total 5 -> 14. Catalog billing blocks: 7 authoritative, 12 curated, 2 derived. 111 tests, ruff and mypy clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Vendors like Speechmatics and Cartesia publish prices only as a web page. Until now that meant I read the page and typed the number into a YAML file, and the `checked_on` date advanced only when someone remembered to look. This adds adapters that read those pages directly, so bud-connect's existing 24h `cron-tensorzero-sync` carries a price change into the catalog on its own. No new scheduler, and nothing in CI: bud-connect already re-seeds every 24 hours via a dapr cron binding, verified live in the budconnect-dev install. The SDK just provides the data; the refresh was already solved. WHAT MAKES THIS SAFE WITHOUT REVIEW There is no human between a page changing and a price being served, so every failure path is designed to end in *the vendor contributing nothing*, never in a wrong number and never in a missing one: page unreachable / non-200 / slow -> vendor skipped adapter raises -> vendor skipped, siblings unaffected fewer models than the adapter's floor -> nothing taken from that vendor a single rate fails a gate -> that model dropped, the rest kept source overruns its 20s budget -> whatever finished is used base catalog empty -> overlay skipped entirely In each case the committed rate in curated_voice_pricing.yaml stands, so scraping can only ever make the catalog staler -- not wrong, not emptier. That matters because bud-connect's seeder reads "absent from this run" as "deactivate", which makes a partial scrape more damaging than no scrape. THE GATES, which are the only guard left once review is gone Bounds per unit (per-second 1e-6..1e-2, per-character 1e-7..1e-3) catch the failures scraping actually produces: a per-hour figure stored as per-second is 3600x out, a misplaced decimal is 10x or 100x. Non-finite rates are rejected explicitly because NaN compares False against every bound and would otherwise pass every range check. `True` is rejected as a rate because it is an int in Python and would pass `> 0`. Duplicate keys reject the whole extraction, since which rate wins would be an accident of ordering and one duplicated row makes the rest of the parse suspect. Above that, the overlay rejects a rate that has moved more than 10x from the committed floor. Vendors halve and double prices -- Speechmatics currently shows a 46% promotional cut -- but they do not move by factors of ten, while a parser reading per-hour as per-second does. An authoritative cloud price is never overwritten by a marketing page, and an entry a live feed priced on a different basis is left alone rather than left with two cost fields disagreeing. PRECEDENCE merge(litellm, ai_models) -> curated floor -> scraped -> authoritative cloud Scraped beats the committed floor, so a new price is taken with no review. A vendor's own price API still beats any page. ADAPTERS speechmatics (7 models) reads the pricing table the page embeds as JSON, taking LIST prices and recording the displayed promotion in the note rather than using it. Bolt-ons (Translation, Summaries, Chapters, Sentiment, Topics) are excluded as per-feature surcharges, like AWS's Call Analytics SKUs. cartesia (2 models, `derived`) parses all three stated conversions -- "1 credit equals 1 character", "3 credits equals 1 second of audio", "at pro tier, it's $65 per 1M credits" -- and fails if any is missing. Nothing is hardcoded, because a hardcoded $65/1M would keep producing plausible numbers long after Cartesia changed it, which is exactly how a derived rate goes quietly wrong. Both were checked against the live pages: 9 prices, 0 rejected, every rate identical to the value I had transcribed by hand. TESTS: 99 new, 210 total Fixtures are trimmed regions of the real pages (27KB and 1.4KB, from 577KB and 192KB) and reproduce the full-page extraction exactly, promotional figures included. No test touches the network -- an adapter that needed a live page to be tested would fail during a vendor's outage. One test asserts each adapter still agrees with the committed floor, because that floor is the fallback used when a scrape fails, and a silent divergence would make the fallback wrong in precisely the situation it exists for. Also: pages are fetched with an honest `bud-model-catalog/1.0` user agent, verified to return identical content to a browser UA, so there is no reason to impersonate one. Concurrency is bounded so eight vendors are not hit simultaneously every 24 hours. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…arges Adds the two remaining adapters for vendors already in the curated file, so all 14 hand-curated prices are now self-refreshing rather than dated by whenever someone last looked. `python -m bud_model_catalog.scrapers --diff` reports 14 prices from 4 vendors with no divergence from the committed file at all. A DATA-LOSS BUG THIS FOUND Rev AI publishes "Rounded up to the nearest second, 15 second minimum" on every model, and the curated entry carried `min_billable_units: 15`. The overlay replaces an entry's whole billing block, so scraping Rev AI would have written a block without that field and silently dropped it -- leaving a rate that under-bills every clip shorter than 15 seconds, with nothing in a diff to show it had happened. Short clips are the common case for voice, so this would have been a real and invisible under-billing. `min_billable_units` now travels the whole path: extracted by the adapter, validated (positive, finite, under an hour -- a "minimum" of thousands of seconds is a parsed year, not a billing floor), carried into the entry, and asserted by a regression test. Absent means the key is omitted rather than set to null, because null reads as "charges from zero", which is a different claim from "the vendor publishes no minimum". REV AI IS PARSED BY STRUCTURE, NOT BY BLOB The first version regexed a tag-stripped page and produced a model named "supports-all-popular-media-types-email-and-chat-support-get-started-view-more-offerings-reverb", because stripping tags joins every element into one line and a non-greedy match had no boundary to stop at. It now walks the page's own `payment-offering` blocks, so a name cannot absorb the sentence in front of it. There is a test pinning that. Excluded: Human Transcription ($1.99/min, which is people), and Forced Alignment, Language Identification, Translation, Sentiment, Summarization and Topic Extraction, which are applied to a transcript rather than deployed -- the same category as AWS's Call Analytics SKUs and Speechmatics' bolt-ons. Rev AI's "per 10 words" units are skipped rather than mis-stored, since the catalog cannot express them. GLADIA TAKES THE LIST PRICE Starter is "Async at $0.61 /hr"; Growth is "as low as $0.20 /hr" and needs an upfront commitment, making it a volume discount rather than a list price. The pattern anchors on "at $" so it cannot pick up a committed rate by accident, and the fixture keeps the Growth card specifically so a test can prove it does not. 226 tests, ruff and mypy clean. Fixtures total 43KB against 1.1MB of original pages. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…wo cannot be
Of the four remaining vendors I said were "static and scrapable", two hold up and two
do not. That earlier classification counted dollar amounts in the HTML without
checking whether they were per-unit rates for the vendor's own models, and it was too
optimistic.
ADDED
google_speech (7 models) is the valuable one: four capabilities and zero models in the
catalog until now, and the only vendor in the set that quotes the catalog's own unit
directly -- "US$0.00003 per character", no conversion to get wrong. Chirp 3 HD $30/1M,
instant custom voice $60/1M, WaveNet and Standard $4/1M, Studio $160/1M, Neural2 and
Polyglot $16/1M. Parsed by table row rather than from stripped text, for the reason
the Rev AI adapter records.
Gemini TTS rows are skipped: they bill per million input text tokens and output audio
tokens, which is a different basis, and storing one as a character rate would be
wrong by orders of magnitude in a way the bounds check would NOT catch. The fixture
keeps a Gemini row so a test holds that line.
Google's speech-to-TEXT page is deliberately not covered. Its table interleaves list
prices with 1-year and 3-year committed-savings columns, so reading a rate means
knowing which column you are in, and getting it wrong under-bills by 20% with nothing
to show it. That needs an adapter written against the structure, not a guess.
speechify (1 entry) charges per character across all voices at a per-plan rate, naming
no priced model, so this is a single `text-to-speech` entry -- a stand-in key following
the label Speechmatics' own table uses, NOT a deployable model name. Worth knowing
before it is shown to anyone as a model. $10/1M entry plan is stored as list; $8 Pro
and $6 Scale are volume tiers.
Its page carries the nastiest trap in this whole exercise: a third-party comparison
of COMPETITORS' prices per million characters, including Cartesia at $49/1M. A
regex for "$X per million characters" finds all of them. Reading Cartesia's rate off
Speechify's marketing page would be absurd on its own, and a future adapter matching
on vendor names could have overwritten the rate this SDK derives from Cartesia's own
page. The pattern anchors on the plan-overage phrasing instead, and a test feeds the
comparison table in and asserts nothing comes out.
NOT ADDED, with the evidence recorded in the YAML header
elevenlabs: 710KB of HTML holding the pricing table's LABELS
("price_per_1k_characters": "Price per 1K characters") and none of its values, which
render client-side. What it does state is credits, marked "approximate", with per-model
variation given as a range ("between 0.5 and 1 credit per character") and no
credit-to-dollar conversion anywhere. Even with a browser that is a range, not a rate.
smallest: prices voice-agent layers approximately ("Text to Speech Layer
~$0.09/minutes"). The only per-model figures on the page belong to third-party LLMs, so
extracting them would attribute OpenAI's prices to smallest. Its own models are never
priced by name.
TWO BUGS FOUND WHILE DOING IT
YAML reads `3e-05` as a string, not a float -- scientific notation needs a decimal
point and signed exponent (`3.0e-05`) to parse as a number. Every generated entry for
the two new vendors arrived as a string, and the failure surfaced as a TypeError
comparing str to int in an unrelated test. Rates are now written in plain decimal, and
CuratedSource rejects a non-numeric rate with a message that says what to do, so the
next person to hand-write a small number gets told rather than debugging a stack trace.
Two scrapers registered for one vendor could produce the same catalog key and silently
overwrite each other -- `validate_batch` catches duplicates within one extraction but
not across two. Exactly how Google would look with both its pages registered and both
yielding a "standard". The clash is now logged as the registry bug it is and the first
value kept.
22 prices from 6 vendors, no divergence from the committed floor. 250 tests, ruff and
mypy clean. Catalog billing blocks: 7 authoritative, 20 curated, 2 derived.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…0% error
Six models: recognition $0.016/min, dynamic batch $0.003, speech recognition with
data logging $0.016 and without $0.024, medical dictation and conversation $0.078.
google_speech now carries 13 models -- 7 text-to-speech, 6 speech-to-text -- where it
had none two commits ago.
WHY THIS ONE WAITED
Every row on this page has three price columns, and in stripped text they are
indistinguishable:
Recognition (sku:3099-B70F-0949) | Standard
| 0 to 500,000 min $0.016 / 1 minute ... <- list
| 0 to 500,000 min $0.0144 / 1 minute ... <- 1-year commitment
| 0 to 500,000 min $0.0128 / 1 minute ... <- 3-year commitment
Storing $0.0144 instead of $0.016 under-bills by 10% and is undetectable: the number
is plausible, the unit is right, the magnitude is right, and neither the bounds check
nor the drift check would ever flag it. Pattern-matching the page the way the other
six adapters do would have been a coin flip, which is why this was left out of the
previous commit rather than guessed at.
Three defences instead:
* The list column is located by its HEADING, never by position, and a table whose
heading cannot be identified yields nothing. There is deliberately no positional
fallback -- "the third column is usually list" is the entire risk.
* Having taken a rate, the adapter asserts it is not cheaper than the same row's
remaining columns. List is by definition dearer than a rate you commit a year to,
so if it is not, the columns were misread and the table is rejected outright.
* The free tier each cell opens with ("0 minute to 60 minute $0.00 (Free)") is
skipped and the first paid tier taken -- the lowest-volume, undiscounted rate, the
same rule applied to AWS tier 0. Taking the first number in the cell instead would
price these models at zero, which is a request billed as free and looks
unremarkable in a catalog.
Unlike the TTS names, parentheses are preserved: "with data logging" and "without data
logging" differ by 50% in price, and the distinction lives entirely inside the
parenthetical.
FRAMEWORK: MORE THAN ONE ADAPTER PER VENDOR
The registry was a dict keyed by vendor, which cannot express "Google publishes TTS and
STT on separate pages with different table shapes". It is now a tuple, and adapters
carry an optional `name` distinct from `vendor` (`google_speech_tts`,
`google_speech_stt`) used to select one on the command line and to find its fixture.
`vendor` remains the catalog key. The cross-adapter key-collision guard added last
commit now has a real case to guard, and a test asserts the two Google adapters claim
no key in common.
The fixture keeps all four pricing tables WITH their savings columns and free tiers,
because those are precisely what the adapter has to avoid reading -- a fixture without
them would test nothing. One test enumerates every committed-savings rate on the page
and asserts none of them was ever stored; another feeds a table whose columns are
swapped and asserts the invariant fires.
28 prices from 7 adapters, no divergence from the committed floor. 268 tests, ruff and
mypy clean. Catalog billing blocks: 7 authoritative, 26 curated, 2 derived.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Speechify meters per character across all voices at a per-plan rate and names no priced model anywhere on its page. The rate is real; there is nothing to attach it to. `speechify/text-to-speech` was a label this project made up, and it would have appeared in budapp's model list as something nobody can deploy -- the same objection that had fireworks and together_ai withdrawn from the provider list. Removes the adapter, its fixture, its tests and its floor entry, and records the reason in the YAML header alongside the other absences. That note also carries the warning for whoever revisits it: Speechify's page lists third-party estimates of COMPETITORS' rates, Cartesia at $49/1M among them, so a future adapter must not match on "$X per million characters". 27 curated prices across 5 vendors, 6 adapters. 261 tests, ruff and mypy clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… models
DEEPGRAM ADVERTISED TEXT-TO-SPEECH WITH NO TTS MODEL
LiteLLM carries 42 Deepgram entries and every one is audio_transcription, so Aura never
reached the catalog. `scrapers/vendors/deepgram.py` supplies it from deepgram.com/pricing:
aura-2 $0.030 / 1k characters (90 voices)
aura $0.015 / 1k characters (12 voices; the pricing page calls it "Aura-1")
Priced per family, not per voice, because that is how Deepgram bills -- the same
granularity as Google's WaveNet and Studio. Keys are the architecture names Deepgram's own
/v1/models uses, not the marketing names, so `aura` rather than `aura-1`. Pay As You Go is
list; Growth is recorded in the note.
Flux TTS is priced on the same page and deliberately NOT emitted. It is absent from
/v1/models, and the only `flux-tts` string in Deepgram's docs is a sidebar URL slug whose
pages return 404, so nothing establishes what `model=` value selects it. A key invented here
would reach budapp as a model nobody can deploy, which is why Speechify was removed.
The curated floor now carries Deepgram. The test that forbade curating a feed-covered
vendor was vendor-level, correct while coverage was all-or-nothing; it is now
(vendor, mode) -- a fed vendor may be curated only in a mode the feed does not supply --
which keeps the intent while letting a genuine gap be filled. If LiteLLM adds Aura later,
the runtime overlay lets the feed win and logs that the curated entry should go.
FIFTEEN ENTRIES THAT WERE NEVER MODELS
`not_a_model()` excludes them at the source. Each rule was checked against the live price
map and removed exactly these and nothing else:
8 no mode fireworks-ai-4.1b-to-16b and siblings: size-bucket pricing tiers.
With no mode there is nothing to derive a route from.
1 guardrail bedrock/guardrails: a filtering product applied to other models.
6 deepgram deepgram/streaming/*: Deepgram's /v1/models has no streaming model and
no `nova-3-multilingual` -- streaming is a way of calling the nova-3
models already listed, and multilingual is nova-3 with language=multi.
detect_entities, diarize, keyterm and redact are feature surcharges.
All fifteen reached bud-connect with no endpoint, and budapp offered them as deployable.
Realtime models are deliberately NOT excluded: they are real models, and whether Bud can
route them is the consumer's question, not a reason to erase them.
29 prices from 7 adapters, 0 rejected. 285 tests; ruff and mypy clean.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…osing fields
A review of every provider's catalog data found entries that were wrong in ways no test
caught. Each is fixed here with a test that fails without the fix.
Prices
- Speechmatics Linden 1 was stored at its $0.30/hr promotion. Its JSON row carries its own
"Crossed-out price" of 0.40 ("discounted 25% from a $0.40/hr list rate"); the adapter now
reads that as list, and the floor is corrected.
- The scraped overlay replaced the whole billing block, deleting floor-only rules such as
Rev AI's rounding_increment from all three Rev AI models. It now merges.
- A scraped key nothing else lists was published unchecked: a renamed row at 10x became a
second model beside the stale first. Pages now refresh prices only; new models enter
through the curated file. Slugs keep '.', so "Linden 1.1" is not "linden-11".
- Every adapter kept "the first" of a repeated row itself, which made validate_batch's
duplicate rejection unreachable. Duplicates are now resolved there only: a repeat that
agrees collapses, one that disagrees rejects the vendor.
- httpx's timeout bounds a read, not a request, so a dribbling server held its slot for the
whole budget. Each vendor now has a real deadline.
Model ids (the key is what budapp sends the vendor as the model)
- Cartesia `sonic` was sunset on 2026-06-01 and `ink` was never an id. Keyed now by the live
ids: sonic-3.6, sonic-3.5, ink-2, ink-whisper.
- ElevenLabs scribe_v1/scribe_v1_experimental and AssemblyAI best/nano are withdrawn from the
vendors' APIs; dropped, with the evidence recorded.
- azure/container, openai/container (Code Interpreter sessions) and Together's size-bucket
tiers were published as models; excluded.
Coverage
- LiteLLM's `vertex_ai` and `vertex_ai-embedding-models` providers were unmapped, so 16 Gemini
models on Vertex (TTS, transcription, Live, embeddings, gemini-omni) never reached the
catalog. All three Vertex providers now feed vertex_ai-gemini-models, Gemini only.
- Colliding LiteLLM keys kept one entry wholesale and discarded the other's fields -- the
Vertex image models lost their modalities and search price. The canonical entry still
wins; the other now fills its gaps.
- New scrapers: ElevenLabs (Flash, which LiteLLM lacks, plus Multilingual, v3, Scribe v2 and
v2 Medical), Deepgram pre-recorded STT (Nova-3; Whisper Large, which LiteLLM had 25% high),
and AssemblyAI (universal-3-5-pro, universal-2, replacing the retired ids).
Live run: 1550 models, 42 prices scraped from 10/10 vendors, 0 rejected.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dittops
marked this pull request as ready for review
September 24, 2026 13:38
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
This PR adds prices for the voice vendors bud-connect already offers. It does not add vendors: offering one Bud cannot dispatch is the mistake the
fireworks/together_aiwithdrawals were about.No aggregator carries dedicated speech vendors. I checked models.dev, OpenRouter and Helicone, and all three share LiteLLM's blind spot. So the prices come from three sources the SDK now layers over LiteLLM, weakest first:
billingblock: the unit, the meter,confidence(authoritative/curated/derived), and the source URL and date. bud-connect stores that block inmodel_info.billing.Pairs with BudEcosystem/bud-connect#37, which consumes it. Merge this first. bud-connect pins
583a24e; if this is squash-merged, bump that pin to the merge commit before #37 merges.1. Authoritative cloud pricing (AWS + Azure)
These feeds are public, unauthenticated and machine-readable. They don't create entries: they attach a
billingblock (region, feed publication date,confidence: authoritative) to entries LiteLLM already has, and check LiteLLM's rate against the vendor's.aws_polly/standard/neural/generative/long-formaws_transcribe/StartTranscriptionJobazure/speech/azure-sttazure/speech/azure-ttsMost of this code exists to avoid four silent failures:
min()would bill everyone the top discount. The code selectsbeginRange == 0instead.+in the query string is read as a space, and the server then returns 200 with an emptyItemsarray. The filter is now percent-encoded, and an empty feed logs a warning.GCP is out of scope because its Billing Catalog refuses unauthenticated callers.
2. Curated floor: 36 entries in
data/curated_voice_pricing.yamlgladia 2 · revai 3 · speechmatics 7 · cartesia 4 · google_speech 13 · deepgram 2 · elevenlabs 3 · assemblyai 2.
source.derived: the page's own conversions multiplied by the Pro-tier credit price, never an invented number.min_billable_units/rounding_increment.FEED_GAPStest).3. Scraped pricing: pages that refresh themselves
bud-connect re-seeds every 24h, so a price change on a vendor's page reaches the catalog within a day with nobody in the loop. With no review step, every failure path has to end in "this vendor contributes nothing and the floor stands", never in a wrong number.
Crossed-out priceis read as list, so Linden 1 is $0.40/hr, not the $0.30 promotion.sonic-3.6,sonic-3.5,ink-2,ink-whisperasync,real-timeaura-2,auranova-3,nova-3-general,whisper-largeeleven_v3,eleven_multilingual_v2,eleven_flash_v2_5,eleven_flash_v2,scribe_v2,scribe_v2_medicaluniversal-3-5-pro,universal-2Validation gates, each with a test:
min_modelsfloor turns a partial extraction into a failure.validate_batchonly, so adapters return every match.Driver,
sources/scraped.py:Deliberately not scraped:
The reasons are written into the YAML header.
4. Catalog correctness fixes found along the way
mode(pricing tiers),bedrock/guardrails,deepgram/streaming/*(pricing SKUs),azure/containerandopenai/container(Code Interpreter sessions), and Together's size-bucket tiers.scribe_v1*and AssemblyAIbest/nano, recorded with evidence inRETIRED_BY_VENDOR.vertex_aiandvertex_ai-embedding-modelsproviders were unmapped, so 16 Gemini-on-Vertex models (TTS, transcription, Live, embeddings) never reached the catalog. All three Vertex providers now feedvertex_ai-gemini-models, filtered to Gemini.sonic/inkto its live ids.Decisions a reviewer should know
Testing
ruff check,ruff format --checkandmypyare clean.budconnect-devand synced into budapp: identical model sets and prices.🤖 Generated with Claude Code