Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Azure OpenAI GPT Realtime API (Microsoft Foundry) Microsoft | OpenAI Realtime API OpenAI | xAI Grok Voice Agent API xAI | |
|---|---|---|---|
| Category | Voice-to-voice | Voice-to-voice | Voice-to-voice |
| Status | GA | GA | GA |
| Est. per minute | $0.096 - 0.84 | $0.096 - 0.76 | $0.08 |
| How that was worked out | Low = gpt-realtime-2.1 Global, 1 min user audio (600 tokens) + 1 min model audio (1,200 tokens), single turn, no cache. Data Zone single turn $0.106. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. High shown at Data Zone rates ($0.76 Global); with full cache hits about $0.06/min. | Low = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached). | Flat $0.08 per minute on grok-voice-think-fast-2.0 regardless of context length. Tool calls and text inputs are extra. xAI does not state whether minutes are session wall-clock time or audio time, so assume wall-clock time including silence. |
| Pricing model | per-token | per-token | per-minute |
| Free tier | No free tier for realtime models (Azure free account credits can apply). | None. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow. | None documented for the voice agent. |
| Connects by | WebRTC, WebSocket, SIP | WebRTC, WebSocket, SIP | WebSocket |
| Audio in | PCM16 mono 24 kHz recommended (send ~100 ms chunks); G.711 supported per the shared OpenAI event model; WebRTC negotiates codecs. | audio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk. | audio/pcm at 8000, 16000, 22050, 24000 (default), 32000, 44100 or 48000 Hz; audio/pcmu, audio/pcma (G.711); audio/opus. JSON base64 or binary transport. |
| Audio out | PCM16 24 kHz (same options as OpenAI). | audio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response. | Same format options as input; speed 0.7 to 1.5. |
| Languages | Same models as OpenAI; Microsoft advises validating languages with production-like audio and passing ISO-639-1 hints for transcription. | Multilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents. | Docs say every voice can speak every supported language; TTS lists 20 languages and STT 38+. No explicit list for the voice agent. |
| Latency (vendor claim) | Microsoft guidance (transport only, not model time): WebRTC ~100 ms, WebSocket ~200 ms. | Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published. | Vendor claim: sub-second latency. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Covered by Azure OpenAI enterprise terms (Microsoft Products and Services DPA; HIPAA BAA via Microsoft for in-scope Azure services). Data Zone deployments keep processing within the US or EU zone. Content filtering applies. | /v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement. | xAI states SOC 2 Type II, HIPAA eligible with a BAA, GDPR with EU data residency options, and that audio is never stored or used for training (vendor claims). |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Flat $/min | - | - | $0.08 |
| Audio in $/1M tok | $32 | $32 | - |
| Audio out $/1M tok | $64 | $64 | - |
| Free tier | No | No | No |
| Free credit $ | - | - | - |
| Native S2S | Yes | Yes | Yes |
| Tools | Yes | Yes | Yes |
| Image in | Yes | Yes | - |
| Own LLM | No | No | No |
| Voices | 10 | 10 | - |
| Cloning | No | No | Yes |
| Latency ms | - | - | - |
| Languages | - | - | - |
| Context tokens | 128,000 | 128,000 | - |
| Max session min | 60 | 60 | - |
| Concurrency | - | - | - |
| WebRTC | Yes | Yes | No |
| WebSocket | Yes | Yes | Yes |
| Phone / SIP | Yes | Yes | No |
| HIPAA | Yes | Yes | Yes |
| SOC 2 | - | Yes | Yes |
| EU data | Yes | Yes | Yes |
| Open weights | No | No | No |
| High warnings | 2 | 3 | 1 |