Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Azure Speech Voice Live API Microsoft | Gemini Live API (Gemini Developer API / Google AI Studio) Google | Gemini Live API on Vertex AI (Gemini Enterprise Agent Platform) Google Cloud | |
|---|---|---|---|
| Category | Voice-to-voice | Voice-to-voice | Voice-to-voice |
| Status | GA | GA | GA |
| Est. per minute | $0.017 - 0.76 | $0.022 - 0.12 | $0.022 - 0.12 |
| How that was worked out | Low = Lite tier phi4-mm-realtime, 1 min user audio (750 tokens x $4/1M) + 1 min model audio (1,200 tokens x $12/1M), single turn. Pro tier with gpt-realtime-2.1 native audio matches Azure OpenAI Global: $0.096 single turn. High = Pro gpt-realtime-2.1, High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Standard tier gpt-realtime-2.1-mini: $0.033 low, $0.26 high. Avatar minutes and custom voice are extra. | Low = gemini-3.8-live, 1 min user audio (1,500 tokens x $3/1M) + 1 min model audio (1,500 tokens x $12/1M), single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Live models do not support context caching, so there is no cached discount. Video input or thinking tokens add more. | Low = 1 min user audio + 1 min model audio at $3 / $12 per 1M, single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Google states tokens from past turns are re-processed and billed every turn up to the context window limit. Live avatar video adds about $0.37 per minute of avatar speech. |
| Pricing model | per-token | per-token | per-token |
| Free tier | No dedicated free tier found for Voice Live (Azure Speech F0 is limited to one concurrent request). | Yes: free tier for all Live models with lower rate limits; free-tier content may be used to improve Google products. | No Live-specific free tier; standard Google Cloud new-customer credits apply. |
| Connects by | WebSocket, WebRTC | WebSocket | WebSocket |
| Audio in | PCM16 mono at 24 kHz (default) or 16 kHz via input_audio_sampling_rate; Live-Reference AEC mode takes interleaved stereo PCM16 (mic + playback reference). Fixed for the session. | Raw 16-bit PCM little-endian, natively 16 kHz (other rates resampled if the MIME type says so, e.g. audio/pcm;rate=16000). Images/video as JPEG or PNG frames, max 1 fps. | Raw 16-bit PCM 16 kHz little-endian; JPEG images/video at 1 fps; text. |
| Audio out | PCM16 audio from the native model or Azure TTS; word timestamps and visemes available with Azure voices; avatar video over WebRTC (H.264). | Raw 16-bit PCM little-endian at 24 kHz (always). | Raw 16-bit PCM 24 kHz little-endian; text; mp4 video for live avatars. |
| Languages | 146 input locales and 151 output locales per the FAQ (Azure speech); azure_semantic_vad_multilingual covers English, Spanish, French, Italian, German, Japanese, Portuguese, Chinese, Korean, Hindi. | Live guide lists 99 languages (overview page says 70); native audio models pick the language automatically and do not accept a language code. | Vertex overview states 24 languages for multilingual support (Developer API docs claim more); verify per language. |
| Latency (vendor claim) | No numeric vendor claim found; marketed as low-latency. Text-LLM cascades are described by Microsoft as having slightly higher latency than native realtime models. | No numeric vendor claim found. | No numeric vendor claim found. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Azure AI services enterprise terms; content filtering always on; data zone and regional model variants for residency. HIPAA BAA via Microsoft for in-scope Azure services (confirm Voice Live scope with Microsoft). | Paid tier data is not used to improve products; free tier is. No BAA or data residency on the Developer API; use Vertex AI for enterprise compliance (CMEK, VPC-SC, regional processing). | Google Cloud terms; CMEK in us/eu multi-regions; customer data not used for training under Google Cloud terms. HIPAA BAA available for covered Google Cloud services (confirm Live API coverage). |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Flat $/min | - | - | - |
| Audio in $/1M tok | $32 | $3 | $3 |
| Audio out $/1M tok | $64 | $12 | $12 |
| Free tier | No | Yes | No |
| Free credit $ | - | - | - |
| Native S2S | Yes | Yes | Yes |
| Tools | Yes | Yes | Yes |
| Image in | Yes | Yes | Yes |
| Own LLM | Yes | No | No |
| Voices | 600 | 30 | 30 |
| Cloning | Yes | - | - |
| Latency ms | - | - | - |
| Languages | 146 | 99 | 24 |
| Context tokens | - | 131,072 | 128,000 |
| Max session min | - | 10 | 10 |
| Concurrency | - | - | 1,000 |
| WebRTC | Yes | No | No |
| WebSocket | Yes | Yes | Yes |
| Phone / SIP | No | No | No |
| HIPAA | Yes | No | Yes |
| SOC 2 | - | - | - |
| EU data | Yes | No | Yes |
| Open weights | No | No | No |
| High warnings | 1 | 3 | 2 |