Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Rev AI Streaming
Rev
Speechmatics Realtime and Agent STT
Speechmatics
OpenAI Realtime transcription sessions
OpenAI
CategorySpeech-to-textSpeech-to-textSpeech-to-text
StatusGAGAGA
Est. per minute$0.0033 - 0.005$0.0033 - 0.013$0.003 - 0.017
How that was worked outAssumes the $0.20-$0.30/hr Reverb lines apply to streaming, which the pricing page does not confirm.Linden 1 with training discount ($0.20/hr) up to Enhanced list ($0.80/hr). Subscriptions/credit packs cut up to 25%/20%.$0.003/min (4o-mini-transcribe estimate, per turn) to $0.017/min (true streaming deltas with gpt-live-transcribe or gpt-realtime-whisper).
Pricing modelper-hourper-hourper-minute
Free tierFree credits equal to 5 hours of Reverb ASR on Pay As You Go.$100 in credits at sign-up, no card (pricing page).None specific to transcription.
Connects byWebSocketWebSocketWebSocket, WebRTC
Audio incontent_type query param, e.g. audio/x-raw;layout=interleaved;rate=16000;format=S16LE;channels=1. Raw, FLAC or WAV recommended.raw pcm_f32le, pcm_s16le or mulaw with explicit sample_rate; or 'file' type for wav, mp3, aac, ogg, mpeg, amr, m4a, mp4, flac.audio/pcm at 24 kHz in the documented example (16-bit, base64 in input_audio_buffer.append); G.711 formats historically supported on Realtime (not re-verified).
Audio outn/an/an/a
Languageslanguage param defaults to en; streaming language list not shown on the API page.Language packs per session (vendor markets 50+ languages; count not re-verified). Enhanced/Standard need a selected language; Melia 1 handles code-switching.Multilingual (count not published on the pages read); gpt-live-transcribe accepts a 'languages' hint list.
Latency (vendor claim)No vendor figure captured.Configurable via max_delay (docs examples use 0.7 s); no other vendor latency figure captured.Vendor: 'low-latency' with tunable delay setting; no number published.
Key limits
  • 3 hours per stream; open a new connection before the limit.
  • Default streaming concurrency 10 (support can raise).
  • max_connection_wait_seconds default 60.
  • Concurrent realtime sessions: Free 2, Pro 50, Enterprise custom.
  • Session ends at 48 hours, after 1 hour without AddAudio, or after 3 minutes with no audio or ping/pong.
  • Custom dictionary over 20,000 items closes the socket with protocol_error.
  • Rate limits depend on account usage tier (not captured).
  • Completion events from different turns can arrive out of order; match by item_id.
High-severity warnings
  • Idle connection time is billed
  • None
  • No server VAD on gpt-live-transcribe
  • True streaming is ~3x the cost of competitors
ComplianceRev markets HIPAA support (site navigation lists HIPAA); not re-verified.Docs state Realtime SaaS does not store audio, transcripts or configuration. On-prem containers and Kubernetes deployments documented. Certifications not re-verified.OpenAI API data controls apply (API data not used for training by default per OpenAI policy); BAA availability not re-verified for these models.
Self-hostableNoYesNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/hour$0.2$0.45$1.02
Bills silenceYes--
Free tierYesYesNo
Free credit $-$100-
Live speakers-YesNo
Turn detect-Yes-
KeytermsYesYesYes
PII redact---
PartialsYesYesYes
Mixed langs-Yes-
8 kHz phone-Yes-
Latency ms---
Languages-50-
Max session min1802,880-
Concurrency1050-
WebRTCNoNoYes
WebSocketYesYesYes
gRPCNoNoNo
HIPAAYes--
SOC 2---
EU data-Yes-
Self-hostNoYesNo
Open weightsNoNoNo
High warnings102