Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Deepgram Voice Agent API
Deepgram
OpenAI GPT-Live API
OpenAI
Ultravox Realtime
Ultravox (formerly Fixie.ai)
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minute$0.041 - 0.163$0.05$0.05 - 0.055
How that was worked outOfficial tier rates; BYO tiers add your own LLM/TTS bills on top.Low = voice layer only, $0.05 per minute of session duration with no delegated work. High depends entirely on the backend model, how often the model delegates and tool fees; budget per-minute session cost plus backend token spend measured in your own tests.$0.05 call rate plus $0.005 SIP when using SIP; LLM and TTS included.
Pricing modelper-minuteper-minuteper-minute
Free tier$200 one-time credit for new accounts.Free tier not supported.30 free call minutes; playground calls free.
Connects byWebSocketWebRTC, WebSocket, SIPWebRTC, WebSocket, SIP, Twilio, Telnyx, Plivo, Exotel
Audio inlinear16 default at 16 kHz (other encodings/rates configurable)WebSocket: audio/pcm 24 kHz (default) or 16 kHz mono 16-bit LE, audio/pcmu or audio/pcma 8 kHz; raw bytes base64, even byte length, no container. Format fixed at session start. WebRTC and SIP negotiate codecs (outbound SIP needs Opus + SDES-SRTP).Raw PCM s16le for server WebSocket (sample rate set in the call medium)
Audio outlinear16 (e.g. 24 kHz), optional container; other encodings configurableSame format as input (one setting covers both).PCM for server WebSocket; WebRTC handles codecs automatically
LanguagesEnglish by default; multilingual via Nova 'multi' or flux-general-multi with language hints, and multilingual TTS providers.Not listed on the model page.Multilingual; no list on the pages read.
Latency (vendor claim)Not stated on the pages read; the API reports per-turn latency metrics (time to first token, TTS time to first byte).Vendor claim: improves Full Duplex Bench score by 30 percentage points over gpt-realtime-2.1 (reported via third-party coverage). No millisecond figure published.Vendor says audio-native processing is faster and robust to transcription errors; no figure on the pages read.
Key limits
  • Concurrency: up to 45 WebSocket agent sessions on Pay As You Go, up to 60 on Growth
  • Billing runs for the whole WebSocket connection
  • Concurrent sessions (OpenAI direct): Build 50, Launch 300, Grow 500; Free unsupported
  • Context window 128,000 tokens; above 90 percent usage a replacement engine starts with up to 8,192 tokens of history, so older details may be summarised or dropped
  • Instructions up to 16,384 tokens; startup history up to 128 messages / 8,192 tokens
  • Session ends with reason expired at a duration limit that the docs do not state
  • Pay As You Go: hard cap of 5 concurrent calls
  • maxDuration defaults to 1 hour per call
  • Invoiced monthly in arrears; early invoice at $10 (PAYG) or $100 (Pro) usage thresholds
High-severity warnings
  • Model choice changes the tier
  • Connection time is billed
  • Two bills, not one
  • Billed by wall-clock duration
  • 5-call concurrency on PAYG
ComplianceNot re-verified this session (Deepgram markets SOC 2 and HIPAA readiness; check the trust page)./v1/live/sessions is ZDR eligible with limitations (store forced false, no forking or recording download). Abuse-monitoring logs 30 days. US and EU data residency.Not stated on the pricing or FAQ pages (a trust portal is linked).
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min$0.075$0.05$0.05
Audio in $/1M tok---
Audio out $/1M tok---
Free tierYesNoYes
Free credit $$200--
Native S2SNoYesNo
ToolsYesYesYes
Image in-No-
Own LLMYesYes-
Voices-12-
Cloning--Yes
Latency ms---
Languages---
Context tokens-128,000-
Max session min---
Concurrency45505
WebRTCNoYesYes
WebSocketYesYesYes
Phone / SIPNoYesYes
HIPAA---
SOC 2---
EU data-Yes-
Open weightsNoNoYes
High warnings221