Catalog

Voice AI models and benchmarks

Every model Speko has measured, and the 45 of 52 catalog entries a request can pin today. Accuracy, latency and price as each board published them — nothing on this page is estimated to fill a column.

Provider

64 models

Speech-to-text

19

WER on FLEURS read English, n=50. Lower is better. Measured 2026-07-03 batch, 2026-07-21 stream.

#ModelProviderId
1
Universal-3.5 Pro
AssemblyAI
2.0%2.0%66ms$0.0075
2
stt-rt-v5
Soniox
benchmark only — not routable
7.5%7.3%78ms$0.0020
3
Ink-2
Cartesia
11.0%9.9%102ms$0.0090
4
Nova-3
Deepgram
9.8%12.9%106ms$0.0048
5
Realtime STT-1
not named
benchmark only — not routable
3.3%3.6%139ms$0.0025
6
Pulse
Smallest AI
5.1%8.0%178ms~$0.0050
7
Scribe v2 Realtime
ElevenLabs
not measured3.4%233ms$0.0065
8
Grok STT
xAI
4.8%10.9%305ms$0.0033
9
Gradium ASR
Gradium
8.4%11.7%334ms$0.0104
10
Flux
Deepgram
not measured6.6%406ms$0.0065
11
Qwen3-ASR
Alibaba
2.8%4.0%424ms$0.0054
12
GPT-4o-mini Transcribe
OpenAI
2.7%6.4%460ms$0.0030
13
GPT-4o Transcribe
OpenAI
2.3%5.8%572ms$0.0060
14
Chirp 3
Google
benchmark only — not routable
3.9%7.4%581ms$0.0160
15
Solaria-1
Gladia
5.0%11.4%596ms$0.0125
16
Velma 2
Modulate
benchmark only — not routable
4.4%5.4%1.11s$0.0010
17
GPT Live Transcribe
OpenAI
not measured4.5%1.12s$0.0170
Scribe v2 / historical batch
ElevenLabs
2.9%not measurednot measurednot measured
GPT Transcribe / batch endpoint
OpenAI
2.5%not measurednot measurednot measured

LLM

18

Multi-step tool tasks completed end-to-end, n=30 runs. Higher is better. Measured 2026-08-08.

#ModelProviderId
1
Claude Haiku 4.5
Anthropic
90%532ms2%0%$5.00
2
DeepSeek-V4-Pronot real-time
Together
83%not measured0%0%$3.48
3
gpt-4.1-minifabricates
OpenAI
83%708ms30%22%$1.60
4
gpt-5.6-luna
OpenAI
100%659ms56%0%$1.20
5
gpt-5.6-terra
OpenAI
100%701ms55%0%$12.00
6
gemini-3.5-flash-lite
Google
67%373ms53%8%$2.50
7
Grok-4.3slow
xAI
100%2.15s46%0%$2.50
8
gpt-4.1
OpenAI
77%640ms41%0%$8.00
9
gpt-oss-120bdead air
Cerebras
80%195ms55%16%$0.75
10
kimi-k2.5retired
LiveKit
benchmark only — not routable
73%1.39s0%0%not measured
11
Claude Sonnet 5
Anthropic
73%1.21s25%0%$10.00
12
gpt-5
OpenAI
70%635ms57%0%$10.00
13
gemma-4-31b
Cerebras
73%192ms38%0%$1.49
14
gemma-4-31b-it
LiveKit
benchmark only — not routable
77%not measured45%0%$1.20
15
gpt-5-mini
OpenAI
60%746ms62%0%$2.00
16
gpt-5-nano
OpenAI
23%525ms45%8%$0.40
17
Llama-3.3-70B
Together
50%698ms65%0%$1.04
18
qwen-turbo
Alibaba
33%483ms0%38%$0.20

Text-to-speech

17

Arena Elo from blind A/B votes, field mean 1500. Higher is better.

#ModelProviderId
1
gemini-3.1-flash-tts-preview
Google
1591978ms~$33.3
2
eleven_v3
ElevenLabs
1590481ms$100.0
3
aura-2
Deepgram
1584125ms$30.0
4
sonic-3.5
Cartesia
1574121ms$50.0
5
simba-3.2
Speechify
1573345ms$10.0
6
tts-rt-v1
Soniox
1569362ms~$13.0
7
s2.1-pro
Fish Audio
benchmark only — not routable
~1566185ms$15.0
8
inworld-tts-2
Inworld
1561116ms$25.0
9
lightning_v3.1
Smallest AI
1544173ms$25.0
10
palabra-tts-v1
Palabra
benchmark only — not routable
~152072ms$30.0
11
grok-tts
xAI
1495272ms$15.0
12
default
Gradium
1448244ms$57.8
13
speech-2.8-hd
MiniMax
1431294ms$100.0
14
arcanav3
Rime
1429238ms$40.0
15
gpt-4o-mini-tts
OpenAI
1424691ms~$20.0
16
octave-2
Hume
1377448ms$100.0
17
qwen3-tts-flash
Alibaba
1310472ms$10.0

Speech-to-speech

10

Completion × capability over 6 concierge scenarios, n=3. Higher is better. GET /v1/models publishes no speech-to-speech entries, so these ids name the measured model rather than a router pin you can send.

#ModelProviderId
1
grok-voice-think-fast-2.0
xAI
xai:grok-voice-think-fast-2.0
0.800.67820ms
2
gemini-3.1-flash-live
Google
google:gemini-3.1-flash-live-preview
0.770.781.11s
3
gpt-realtime
OpenAI
openai:gpt-realtime
0.760.83494ms
4
gpt-realtime-2.1-mini
OpenAI
openai:gpt-realtime-2.1-mini
0.750.831.01s
5
grok-voice-fast
xAI
xai:grok-voice-fast-1.0
0.720.441.11s
6
gpt-realtime-2
OpenAI
openai:gpt-realtime-2
0.680.871.10s
7
gpt-realtime-2.1
OpenAI
openai:gpt-realtime-2.1
0.630.561.10s
8
gpt-realtime-mini
OpenAI
openai:gpt-realtime-mini
0.620.70614ms
9
grok-voice-think-fast-1.0
xAI
xai:grok-voice-think-fast-1.0
0.560.221.01s
10
gpt-4.1-nano (cascade)
Inworld
inworld:openai/gpt-4.1-nano
0.451.00not measured

Measured per language

Each language is its own study over its own field, so these tables replace the English boards rather than filtering them — 17 studies across 9 languages, ranking models the English run never included. The model that wins one study routinely loses another, which is the whole argument for routing per language instead of picking one stack and hoping.

Arabic speech-to-text

#ModelCERbatch
1GPT-4o Transcribe2.7%
2Whisper-12.8%
3Qwen3-ASR3.5%
4GPT-4o-mini Transcribe4.7%
5stt-rt-v55.1%
6Nova-37.1%
7Ink-Whisper10.2%

CER / batch. Lower is better. Measured 2026-07-21.

Arabic text-to-speech

#ModelNaturalnessinternal MOS
1grok-tts4.15
2sonic-3.53.73
3eleven_v33.33
4inworld-tts-22.76

Naturalness / internal MOS. Higher is better. Blind human A/B listening, scored as internal MOS. Measured 2026-07-21.

Published stacks

What benchmarks.speko.ai picks for each job off the boards above, and the measurement that decided it. A stack whose legs are identical to another’s is one row: the publisher distinguishes four jobs and its boards distinguish two stacks. A stack with a leg the router cannot currently call is not shown at all — a measurement may stay visible in that state, an instruction you would copy may not.

Use case
Real-time phone agentdecided on lowest-latency legs
Speech-to-textUniversal-3.5 Proassemblyai:universal-3-5-pro
LLMgemma-4-31bcerebras:gemma-4-31b
Text-to-speechinworld-tts-2inworld:inworld-tts-2
Accuracy-criticaldecided on 0% fabricationTool-heavy agentdecided on 3% tool silence
Speech-to-textUniversal-3.5 Proassemblyai:universal-3-5-pro
LLMClaude Haiku 4.5anthropic:claude-haiku-4-5
Text-to-speechinworld-tts-2inworld:inworld-tts-2
Natural conversationdecided on 2% dead-air
Speech-to-textUniversal-3.5 Proassemblyai:universal-3-5-pro
LLMClaude Haiku 4.5anthropic:claude-haiku-4-5
Text-to-speecheleven_v3elevenlabs:eleven_v3

How to read this

An id is a pin. Where a row shows a provider:model string, that string is what you send as model to pin the route; click it to copy it. Where it says benchmark only, the public catalog does not currently offer that model, and the number beside it is evidence rather than an instruction. The boards and the catalog disagree about 7 models today, and the catalog wins every time — it is the surface your key hits.

Not measured is not zero. An empty cell says not measured in words, because a dash in a column of word error rates reads as a perfect score. A leading ~ marks a figure the board itself publishes as an estimate — a vendor with no exact published rate, or an Elo fit over too few votes.

Nothing here re-ranks a board. Tables arrive in the order the publisher authored, and the # column is an index into whatever order you have sorted them into — not a verdict. A study that publishes no positions shows none.