Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Phonic
Phonic
OpenAI Realtime API
OpenAI
Azure OpenAI GPT Realtime API (Microsoft Foundry)
Microsoft
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minute$0.15$0.096 - 0.76$0.096 - 0.84
How that was worked outPublished starting price; volume pricing via sales.Low = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached).Low = gpt-realtime-2.1 Global, 1 min user audio (600 tokens) + 1 min model audio (1,200 tokens), single turn, no cache. Data Zone single turn $0.106. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. High shown at Data Zone rates ($0.76 Global); with full cache hits about $0.06/min.
Pricing modelper-minuteper-tokenper-token
Free tierNot stated.None. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow.No free tier for realtime models (Azure free account credits can apply).
Connects byWebSocket, Webhooks, SIP (Twilio, Telnyx), LiveKit plugin, Amazon Connect via Chime SIPWebRTC, WebSocket, SIPWebRTC, WebSocket, SIP
Audio inpcm_44100 (default), pcm_24000, pcm_16000, pcm_8000, mulaw_8000audio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk.PCM16 mono 24 kHz recommended (send ~100 ms chunks); G.711 supported per the shared OpenAI event model; WebRTC negotiates codecs.
Audio outSame optionsaudio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response.PCM16 24 kHz (same options as OpenAI).
Languages51 languages (per docs index).Multilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents.Same models as OpenAI; Microsoft advises validating languages with production-like audio and passing ISO-639-1 hints for transcription.
Latency (vendor claim)Vendor claim: sub-500 ms speech-in to speech-out.Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published.Microsoft guidance (transport only, not model time): WebRTC ~100 ms, WebSocket ~200 ms.
Key limits
  • Rate limits per organization: 500 requests/second shared across /sts/ws and SIP outbound; 5 requests/second for /conversations/outbound_call
  • Max session length 60 minutes, no warning event before cutoff (third-party SDKs reconnect at ~50 min)
  • Context window 128,000 tokens; max output 32,000 (2.1 and 2.1-mini)
  • Default rate limits gpt-realtime-2.1: Build 400 RPM / 200,000 TPM, Launch 10,000 RPM / 4,000,000 TPM, Grow 20,000 RPM / 15,000,000 TPM
  • gpt-realtime-translate: Build 200, Launch 650, Grow 850 minutes of audio per minute
  • Max session duration 60 minutes (monitor expires_at in session.created)
  • Default gpt-realtime quota in the Tier 1 table: 200 RPM and 100,000 TPM (GlobalStandard); higher tiers raise it (e.g. 300 RPM / 150,000 TPM)
  • GPT-Live on Azure: concurrent sessions per subscription Default 10, Tier 1 25, Tier 2 50, Tier 3 200, Tier 4 300, Tier 5 500
  • Realtime quota is separate from chat completions quota
High-severity warnings
  • None
  • Context growth multiplies cost
  • Beta interface is gone
  • Deprecation calendar
  • Preview endpoints and samples are deprecated
  • Low default TPM
ComplianceHIPAA and SOC 2 compliance, 99.9% uptime SLA (vendor claim)./v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement.Covered by Azure OpenAI enterprise terms (Microsoft Products and Services DPA; HIPAA BAA via Microsoft for in-scope Azure services). Data Zone deployments keep processing within the US or EU zone. Content filtering applies.
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min$0.15--
Audio in $/1M tok-$32$32
Audio out $/1M tok-$64$64
Free tier-NoNo
Free credit $---
Native S2SYesYesYes
ToolsYesYesYes
Image in-YesYes
Own LLMNoNoNo
Voices-1010
Cloning-NoNo
Latency ms500--
Languages51--
Context tokens-128,000128,000
Max session min-6060
Concurrency---
WebRTCNoYesYes
WebSocketYesYesYes
Phone / SIPYesYesYes
HIPAAYesYesYes
SOC 2YesYes-
EU data-YesYes
Open weightsNoNoNo
High warnings032