Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Inworld Realtime API Inworld AI | Azure Speech Voice Live API Microsoft | Alibaba Cloud Model Studio Qwen-Omni-Realtime Alibaba Cloud | |
|---|---|---|---|
| Category | Voice-to-voice | Voice-to-voice | Voice-to-voice |
| Status | GA | GA | GA |
| Est. per minute | $0.01 - 0.03 | $0.017 - 0.76 | $0.0018 - 0.015 |
| How that was worked out | Own estimate: STT about $0.0025/min plus TTS-2 about $0.0125-0.025 per minute of agent speech (1,000 chars/min), before LLM cost. | Low = Lite tier phi4-mm-realtime, 1 min user audio (750 tokens x $4/1M) + 1 min model audio (1,200 tokens x $12/1M), single turn. Pro tier with gpt-realtime-2.1 native audio matches Azure OpenAI Global: $0.096 single turn. High = Pro gpt-realtime-2.1, High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Standard tier gpt-realtime-2.1-mini: $0.033 low, $0.26 high. Avatar minutes and custom voice are extra. | Low = qwen3.8-omni-flash-realtime, 1 min user audio (420 tokens x $0.93/1M) + 1 min model audio (750 tokens x $1.87/1M), single turn, excluding the matching output text tokens (small). High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Alibaba documents that each turn re-bills retained history. qwen3.5-omni-plus-realtime: $0.053 low, $0.27 high. |
| Pricing model | per-character | per-token | per-token |
| Free tier | On-Demand plan free to start with up to 70 TTS minutes and up to 400 STT minutes. | No dedicated free tier found for Voice Live (Azure Speech F0 is limited to one concurrent request). | 1M free tokens per model in the Singapore region for 90 days. |
| Connects by | WebSocket, WebRTC | WebSocket, WebRTC | WebSocket, WebRTC |
| Audio in | PCM16 mono 24 kHz default; G.711 mu-law/A-law 8 kHz; float32 | PCM16 mono at 24 kHz (default) or 16 kHz via input_audio_sampling_rate; Live-Reference AEC mode takes interleaved stereo PCM16 (mic + playback reference). Fixed for the session. | PCM 16 kHz (16-bit mono LE); multichannel 2 or 4 channel spatial input supported (double tokens); JPG images about 1 fps for video. |
| Audio out | Same options | PCM16 audio from the native model or Azure TTS; word timestamps and visemes available with Azure voices; avatar video over WebRTC (H.264). | PCM 24 kHz. |
| Languages | Depends on STT/TTS models chosen; not listed on the pages read. | 146 input locales and 151 output locales per the FAQ (Azure speech); azure_semantic_vad_multilingual covers English, Spanish, French, Italian, German, Japanese, Portuguese, Chinese, Korean, Hindi. | Speech recognition 113 languages and dialects; speech generation 36 languages and dialects. |
| Latency (vendor claim) | Not stated on the pages read. | No numeric vendor claim found; marketed as low-latency. Text-LLM cascades are described by Microsoft as having slightly higher latency than native realtime models. | No numeric vendor claim captured. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Not stated on the pages read. | Azure AI services enterprise terms; content filtering always on; data zone and regional model variants for residency. HIPAA BAA via Microsoft for in-scope Azure services (confirm Voice Live scope with Microsoft). | Data retention, residency guarantees and certifications for Model Studio realtime were not captured; review Alibaba Cloud International terms before sending regulated data. |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Flat $/min | - | - | - |
| Audio in $/1M tok | - | $32 | $0.93 |
| Audio out $/1M tok | - | $64 | $1.87 |
| Free tier | Yes | No | Yes |
| Free credit $ | - | - | - |
| Native S2S | No | Yes | Yes |
| Tools | Yes | Yes | Yes |
| Image in | - | Yes | Yes |
| Own LLM | - | Yes | No |
| Voices | - | 600 | - |
| Cloning | Yes | Yes | Yes |
| Latency ms | - | - | - |
| Languages | - | 146 | 36 |
| Context tokens | - | - | 196,608 |
| Max session min | - | - | 120 |
| Concurrency | 5 | - | - |
| WebRTC | Yes | Yes | Yes |
| WebSocket | Yes | Yes | Yes |
| Phone / SIP | No | No | No |
| HIPAA | - | Yes | - |
| SOC 2 | - | - | - |
| EU data | - | Yes | No |
| Open weights | No | No | No |
| High warnings | 1 | 1 | 2 |