Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

OpenAI Realtime transcription sessions
OpenAI
Speechmatics Realtime and Agent STT
Speechmatics
Rev AI Streaming
Rev
CategorySpeech-to-textSpeech-to-textSpeech-to-text
StatusGAGAGA
Est. per minute$0.003 - 0.017$0.0033 - 0.013$0.0033 - 0.005
How that was worked out$0.003/min (4o-mini-transcribe estimate, per turn) to $0.017/min (true streaming deltas with gpt-live-transcribe or gpt-realtime-whisper).Linden 1 with training discount ($0.20/hr) up to Enhanced list ($0.80/hr). Subscriptions/credit packs cut up to 25%/20%.Assumes the $0.20-$0.30/hr Reverb lines apply to streaming, which the pricing page does not confirm.
Pricing modelper-minuteper-hourper-hour
Free tierNone specific to transcription.$100 in credits at sign-up, no card (pricing page).Free credits equal to 5 hours of Reverb ASR on Pay As You Go.
Connects byWebSocket, WebRTCWebSocketWebSocket
Audio inaudio/pcm at 24 kHz in the documented example (16-bit, base64 in input_audio_buffer.append); G.711 formats historically supported on Realtime (not re-verified).raw pcm_f32le, pcm_s16le or mulaw with explicit sample_rate; or 'file' type for wav, mp3, aac, ogg, mpeg, amr, m4a, mp4, flac.content_type query param, e.g. audio/x-raw;layout=interleaved;rate=16000;format=S16LE;channels=1. Raw, FLAC or WAV recommended.
Audio outn/an/an/a
LanguagesMultilingual (count not published on the pages read); gpt-live-transcribe accepts a 'languages' hint list.Language packs per session (vendor markets 50+ languages; count not re-verified). Enhanced/Standard need a selected language; Melia 1 handles code-switching.language param defaults to en; streaming language list not shown on the API page.
Latency (vendor claim)Vendor: 'low-latency' with tunable delay setting; no number published.Configurable via max_delay (docs examples use 0.7 s); no other vendor latency figure captured.No vendor figure captured.
Key limits
  • Rate limits depend on account usage tier (not captured).
  • Completion events from different turns can arrive out of order; match by item_id.
  • Concurrent realtime sessions: Free 2, Pro 50, Enterprise custom.
  • Session ends at 48 hours, after 1 hour without AddAudio, or after 3 minutes with no audio or ping/pong.
  • Custom dictionary over 20,000 items closes the socket with protocol_error.
  • 3 hours per stream; open a new connection before the limit.
  • Default streaming concurrency 10 (support can raise).
  • max_connection_wait_seconds default 60.
High-severity warnings
  • No server VAD on gpt-live-transcribe
  • True streaming is ~3x the cost of competitors
  • None
  • Idle connection time is billed
ComplianceOpenAI API data controls apply (API data not used for training by default per OpenAI policy); BAA availability not re-verified for these models.Docs state Realtime SaaS does not store audio, transcripts or configuration. On-prem containers and Kubernetes deployments documented. Certifications not re-verified.Rev markets HIPAA support (site navigation lists HIPAA); not re-verified.
Self-hostableNoYesNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/hour$1.02$0.45$0.2
Bills silence--Yes
Free tierNoYesYes
Free credit $-$100-
Live speakersNoYes-
Turn detect-Yes-
KeytermsYesYesYes
PII redact---
PartialsYesYesYes
Mixed langs-Yes-
8 kHz phone-Yes-
Latency ms---
Languages-50-
Max session min-2,880180
Concurrency-5010
WebRTCYesNoNo
WebSocketYesYesYes
gRPCNoNoNo
HIPAA--Yes
SOC 2---
EU data-Yes-
Self-hostNoYesNo
Open weightsNoNoNo
High warnings201