Best AI Speech-to-text models in 2026
Transcription, captions and voice-agent input.
Budget pick: Qwen3-ASR, Whisper

AssemblyAI Universal
AssemblyAI · Universal-3.5 Pro
APIAccurate batch transcription
Topped Coval's 31-model STT benchmark in July 2026 at 3.3% word error rate. Streaming version for voice agents.
- Where to use:
- AssemblyAI API ($50 free credit on signup).

ElevenLabs Scribe
ElevenLabs · Scribe v2
APIMultilingual transcripts with speaker labels
99 languages, speaker diarization and audio-event tags like laughter. Has a real-time variant.
- Where to use:
- ElevenLabs app and API.

Deepgram Nova
Deepgram · Nova-3
APILow-latency streaming at scale
Fast, cheap streaming transcription widely used in call centers and voice agents.
- Where to use:
- Deepgram API ($200 free credit, no card).

Qwen3-ASR
Alibaba · Qwen3-ASR Flash
APIBudget pickVery cheap, accurate, Chinese dialects
Matched AssemblyAI in a human-expert accuracy test at a fraction of the price. 11 languages plus Chinese dialects, robust to noise, and accepts context text to bias vocabulary.
- Where to use:
- Alibaba Cloud Model Studio, OpenRouter.
- Price:
- ≈$0.13 / hour of audio

Whisper
OpenAI · Whisper large-v3
Open weightsBudget pickFree, open-weight transcription
MIT-licensed and runs locally, from laptops (whisper.cpp) to servers. The base for countless transcription apps.
- Where to use:
- GitHub, Hugging Face, or hosted on Groq and others.
New model versions ship every few weeks and rankings shift with them. Prices are rough API list rates and vary by host and resolution. Check the official page before you commit.