Speech-to-text, live

Streaming transcription with partial results and turn detection, for captions, notes and voice agents.

Columns 10/31
Basics
Cost
Features
Performance
Limits
Connects by
Compliance
Openness
Warnings
Compare 0
Must have:

Showing 19 of 19. Click any column header to sort.

Groq Whisper (no streaming)Groq
GA$0.00067$0.0019$0.04YesNoNo--1
SonioxSoniox
GA$0.002$0.0025$0.12-YesYes-601
Cartesia Ink (ink-2, ink-whisper)Cartesia
GA$0.0022$0.012$0.39Yes-Yes--1
AssemblyAIAssemblyAI
GA$0.0025$0.0095$0.45YesYesYes-322
OpenAI transcriptionOpenAI
GA$0.003$0.017$1.02NoNo---2
Rev AI StreamingRev
GA$0.0033$0.005$0.2Yes----1
Speechmatics Realtime and Agent STTSpeechmatics
GA$0.0033$0.013$0.45YesYesYes-500
Google Cloud STTGoogle Cloud
GA$0.004$0.016$0.96YesNo---2
Smallest.ai Pulse STTSmallest AI
GA$0.004$0.006$0.36-YesYes-210
DeepgramDeepgram
GA$0.0042$0.0098$0.288YesYesYes260592
Gladia Live (Solaria-1)Gladia
GA$0.0042$0.013$0.75Yes-Yes-1001
Sarvam AI Saaras Realtime STTSarvam AI
GA$0.0059$0.0088$0.35Yes----0
Voxtral RealtimeMistral AI
GA$0.006$0.006$0.36-No-200130
Voxist ASR (via OVHcloud / Scaleway marketplaces)Voxist
GA$0.0063$0.016--No---0
ElevenLabs ScribeElevenLabs
GA$0.0065$0.0098$0.39Yes-Yes150900
Azure Speech STTMicrosoft
GA$0.0067$0.025$1YesYes---1
Amazon TranscribeAWS
GA$0.01$0.075$0.6YesYes--771
Picovoice Cheetah (on-device streaming STT)Picovoice
GA$0.02$0.02$1.2Yes-Yes-81
Fireworks AI streaming ASR (deprecated)Fireworks AI
Deprecated$0$0------1

A dash means the vendor does not say. Latency is the vendor's own claim, not our measurement. Per-minute figures are estimates from list prices; each API page explains the basis, because vendors bill by tokens, characters, connection time or flat minutes.

Category warnings

These apply to most APIs in this category.

Billing basis differs: socket time vs audio vs tokens

AssemblyAI bills the time the WebSocket is open; Rev AI bills the larger of stream time and audio time with a 15 s minimum; Cartesia bills audio seconds including silence; Soniox and older OpenAI models bill tokens. An idle open connection can cost the same as speech, so close sessions as soon as a call ends.

Idle timeouts and keepalives

Deepgram closes after 10 s with no audio or KeepAlive; Speechmatics after 3 minutes without audio or ping; AssemblyAI has an inactivity_timeout and a 3-hour cap. Mute buttons, hold music gaps and slow LLM turns are the usual triggers. Send keepalives or silence and handle reconnects.

Maximum stream length

Google V2 streaming is about 5 minutes per stream; AssemblyAI, Gladia and Rev AI cap at 3 hours; Soniox 300 minutes (fixed); Speechmatics 48 hours. Long meetings need overlapping reconnects and transcript stitching.

Partial vs final results

Partials (interim, non-final tokens) are rewritten as more audio arrives. Only commit finals to storage, LLM context or analytics; render partials as provisional UI. Some APIs send deltas (Cartesia, OpenAI) that must be concatenated exactly as received.

Endpointing is a latency vs cut-off trade-off

Short silence thresholds make voice agents snappy but split sentences at natural pauses; long thresholds add hundreds of ms per turn. Turn-aware models (Deepgram Flux, AssemblyAI turn detection, Speechmatics Agent STT, Smallest pulse-2, Soniox semantic endpointing) help, but tune per use case and language.

Streaming diarization is often missing or extra

Google Chirp 3 (batch only), Mistral Voxtral Realtime, ElevenLabs Scribe Realtime and Gladia live do not document live speaker labels; Deepgram and AssemblyAI charge extra on streaming. For phone calls, send each party as its own channel instead.

8 kHz telephony audio

Send phone audio at its native 8 kHz mu-law/PCM when the API supports it (Deepgram mulaw, ElevenLabs ulaw_8000, Speechmatics mulaw, Cartesia pcm_mulaw, Google MULAW) instead of upsampling; accuracy on narrowband audio is lower than vendor benchmarks, so test with real calls.

Never ship long-lived keys to browsers

Use short-lived credentials: Deepgram JWT (30 s TTL), AssemblyAI temporary tokens, Speechmatics JWT, ElevenLabs single-use tokens, Mistral rt_ tokens (~900 s), Cartesia access tokens, OpenAI ephemeral client secrets, Gladia per-session URLs. Some APIs (Rev AI, Sarvam) only offer key-in-URL or subprotocol auth, so relay through your server.

Real-time pacing and chunk size

Most APIs expect audio at about real-time speed in 20-200 ms chunks; AssemblyAI rejects chunks outside 50-1000 ms or faster-than-real-time sends, Deepgram Flux wants 80 ms, AWS recommends 50-200 ms. Streaming a file needs a sleep between chunks.

Rapid model churn in 2026

Soniox v4 to v5 (auto-routed), AssemblyAI moved to universal-3-6-pro and dropped u3-rt-pro IDs, OpenAI added gpt-realtime-whisper then gpt-live-transcribe and gpt-transcribe, Fireworks deprecated audio. Pin model IDs, watch changelogs and re-test before upgrading.

Promotional and unlabeled prices

Deepgram shows streaming rates as limited-time promotional, Azure lists MAI-Transcribe-2 under a promotion to 2026-12-31, and several vendors (Rev AI, Sarvam) do not separate streaming from batch prices. Confirm the streaming line item in writing.

Latency claims are not comparable

Vendors measure differently (first partial, end-of-turn, final). Benchmark time-to-final on your own audio, network and region instead of comparing marketing numbers.