Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

OpenAI Realtime API
OpenAI
Azure OpenAI GPT Realtime API (Microsoft Foundry)
Microsoft
xAI Grok Voice Agent API
xAI
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minute$0.096 - 0.76$0.096 - 0.84$0.08
How that was worked outLow = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached).Low = gpt-realtime-2.1 Global, 1 min user audio (600 tokens) + 1 min model audio (1,200 tokens), single turn, no cache. Data Zone single turn $0.106. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. High shown at Data Zone rates ($0.76 Global); with full cache hits about $0.06/min.Flat $0.08 per minute on grok-voice-think-fast-2.0 regardless of context length. Tool calls and text inputs are extra. xAI does not state whether minutes are session wall-clock time or audio time, so assume wall-clock time including silence.
Pricing modelper-tokenper-tokenper-minute
Free tierNone. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow.No free tier for realtime models (Azure free account credits can apply).None documented for the voice agent.
Connects byWebRTC, WebSocket, SIPWebRTC, WebSocket, SIPWebSocket
Audio inaudio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk.PCM16 mono 24 kHz recommended (send ~100 ms chunks); G.711 supported per the shared OpenAI event model; WebRTC negotiates codecs.audio/pcm at 8000, 16000, 22050, 24000 (default), 32000, 44100 or 48000 Hz; audio/pcmu, audio/pcma (G.711); audio/opus. JSON base64 or binary transport.
Audio outaudio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response.PCM16 24 kHz (same options as OpenAI).Same format options as input; speed 0.7 to 1.5.
LanguagesMultilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents.Same models as OpenAI; Microsoft advises validating languages with production-like audio and passing ISO-639-1 hints for transcription.Docs say every voice can speak every supported language; TTS lists 20 languages and STT 38+. No explicit list for the voice agent.
Latency (vendor claim)Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published.Microsoft guidance (transport only, not model time): WebRTC ~100 ms, WebSocket ~200 ms.Vendor claim: sub-second latency.
Key limits
  • Max session length 60 minutes, no warning event before cutoff (third-party SDKs reconnect at ~50 min)
  • Context window 128,000 tokens; max output 32,000 (2.1 and 2.1-mini)
  • Default rate limits gpt-realtime-2.1: Build 400 RPM / 200,000 TPM, Launch 10,000 RPM / 4,000,000 TPM, Grow 20,000 RPM / 15,000,000 TPM
  • gpt-realtime-translate: Build 200, Launch 650, Grow 850 minutes of audio per minute
  • Max session duration 60 minutes (monitor expires_at in session.created)
  • Default gpt-realtime quota in the Tier 1 table: 200 RPM and 100,000 TPM (GlobalStandard); higher tiers raise it (e.g. 300 RPM / 150,000 TPM)
  • GPT-Live on Azure: concurrent sessions per subscription Default 10, Tier 1 25, Tier 2 50, Tier 3 200, Tier 4 300, Tier 5 500
  • Realtime quota is separate from chat completions quota
  • Session duration and concurrency limits are not published
  • Resumption history dropped after 30 minutes of inactivity
  • keyterms: max 100 terms, 50 characters each
  • Ephemeral client secrets: expiry set by expires_after.seconds (example 300); session and anchor fields not supported
High-severity warnings
  • Context growth multiplies cost
  • Beta interface is gone
  • Deprecation calendar
  • Preview endpoints and samples are deprecated
  • Low default TPM
  • Alias moved and price rose
Compliance/v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement.Covered by Azure OpenAI enterprise terms (Microsoft Products and Services DPA; HIPAA BAA via Microsoft for in-scope Azure services). Data Zone deployments keep processing within the US or EU zone. Content filtering applies.xAI states SOC 2 Type II, HIPAA eligible with a BAA, GDPR with EU data residency options, and that audio is never stored or used for training (vendor claims).
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min--$0.08
Audio in $/1M tok$32$32-
Audio out $/1M tok$64$64-
Free tierNoNoNo
Free credit $---
Native S2SYesYesYes
ToolsYesYesYes
Image inYesYes-
Own LLMNoNoNo
Voices1010-
CloningNoNoYes
Latency ms---
Languages---
Context tokens128,000128,000-
Max session min6060-
Concurrency---
WebRTCYesYesNo
WebSocketYesYesYes
Phone / SIPYesYesNo
HIPAAYesYesYes
SOC 2Yes-Yes
EU dataYesYesYes
Open weightsNoNoNo
High warnings321