Catalog
Voice AI models and benchmarks
Every model Speko has measured, and the 45 of 52 catalog entries a request can pin today. Accuracy, latency and price as each board published them — nothing on this page is estimated to fill a column.
64 models
Speech-to-text
19WER on FLEURS read English, n=50. Lower is better. Measured 2026-07-03 batch, 2026-07-21 stream.
LLM
18Multi-step tool tasks completed end-to-end, n=30 runs. Higher is better. Measured 2026-08-08.
Text-to-speech
17Arena Elo from blind A/B votes, field mean 1500. Higher is better.
Speech-to-speech
10Completion × capability over 6 concierge scenarios, n=3. Higher is better. GET /v1/models publishes no speech-to-speech entries, so these ids name the measured model rather than a router pin you can send.
Measured per language
Each language is its own study over its own field, so these tables replace the English boards rather than filtering them — 17 studies across 9 languages, ranking models the English run never included. The model that wins one study routinely loses another, which is the whole argument for routing per language instead of picking one stack and hoping.
Arabic speech-to-text
CER / batch. Lower is better. Measured 2026-07-21.
French speech-to-text
WER / batch. Lower is better. Measured 2026-07-21.
German speech-to-text
WER / batch. Lower is better. Measured 2026-07-21.
Hindi speech-to-text
WER / batch. Lower is better. Measured 2026-07-21.
Norwegian speech-to-text
WER / stream. Lower is better. Measured 2026-07-21.
Spanish speech-to-text
WER / batch. Lower is better. Measured 2026-07-21.
Tamil speech-to-text
CER / batch. Lower is better. Measured 2026-07-21.
Telugu speech-to-text
CER / batch. Lower is better. Measured 2026-07-21.
Arabic text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
Filipino text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21. The board names no winner in this study: the measured intervals overlap.
French text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
German text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
Hindi text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
Norwegian text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
Spanish text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
Tamil text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
Telugu text-to-speech
Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.
Published stacks
What benchmarks.speko.ai picks for each job off the boards above, and the measurement that decided it. A stack whose legs are identical to another’s is one row: the publisher distinguishes four jobs and its boards distinguish two stacks. A stack with a leg the router cannot currently call is not shown at all — a measurement may stay visible in that state, an instruction you would copy may not.
How to read this
An id is a pin. Where a row shows a provider:model string, that string is what you send as model to pin the route; click it to copy it. Where it says benchmark only, the public catalog does not currently offer that model, and the number beside it is evidence rather than an instruction. The boards and the catalog disagree about 7 models today, and the catalog wins every time — it is the surface your key hits.
Not measured is not zero. An empty cell says not measured in words, because a dash in a column of word error rates reads as a perfect score. A leading ~ marks a figure the board itself publishes as an estimate — a vendor with no exact published rate, or an Elo fit over too few votes.
Nothing here re-ranks a board. Tables arrive in the order the publisher authored, and the # column is an index into whatever order you have sorted them into — not a verdict. A study that publishes no positions shows none.