Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Inworld Realtime API
Inworld AI
Azure Speech Voice Live API
Microsoft
Alibaba Cloud Model Studio Qwen-Omni-Realtime
Alibaba Cloud
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minute$0.01 - 0.03$0.017 - 0.76$0.0018 - 0.015
How that was worked outOwn estimate: STT about $0.0025/min plus TTS-2 about $0.0125-0.025 per minute of agent speech (1,000 chars/min), before LLM cost.Low = Lite tier phi4-mm-realtime, 1 min user audio (750 tokens x $4/1M) + 1 min model audio (1,200 tokens x $12/1M), single turn. Pro tier with gpt-realtime-2.1 native audio matches Azure OpenAI Global: $0.096 single turn. High = Pro gpt-realtime-2.1, High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Standard tier gpt-realtime-2.1-mini: $0.033 low, $0.26 high. Avatar minutes and custom voice are extra.Low = qwen3.8-omni-flash-realtime, 1 min user audio (420 tokens x $0.93/1M) + 1 min model audio (750 tokens x $1.87/1M), single turn, excluding the matching output text tokens (small). High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Alibaba documents that each turn re-bills retained history. qwen3.5-omni-plus-realtime: $0.053 low, $0.27 high.
Pricing modelper-characterper-tokenper-token
Free tierOn-Demand plan free to start with up to 70 TTS minutes and up to 400 STT minutes.No dedicated free tier found for Voice Live (Azure Speech F0 is limited to one concurrent request).1M free tokens per model in the Singapore region for 90 days.
Connects byWebSocket, WebRTCWebSocket, WebRTCWebSocket, WebRTC
Audio inPCM16 mono 24 kHz default; G.711 mu-law/A-law 8 kHz; float32PCM16 mono at 24 kHz (default) or 16 kHz via input_audio_sampling_rate; Live-Reference AEC mode takes interleaved stereo PCM16 (mic + playback reference). Fixed for the session.PCM 16 kHz (16-bit mono LE); multichannel 2 or 4 channel spatial input supported (double tokens); JPG images about 1 fps for video.
Audio outSame optionsPCM16 audio from the native model or Azure TTS; word timestamps and visemes available with Azure voices; avatar video over WebRTC (H.264).PCM 24 kHz.
LanguagesDepends on STT/TTS models chosen; not listed on the pages read.146 input locales and 151 output locales per the FAQ (Azure speech); azure_semantic_vad_multilingual covers English, Spanish, French, Italian, German, Japanese, Portuguese, Chinese, Korean, Hindi.Speech recognition 113 languages and dialects; speech generation 36 languages and dialects.
Latency (vendor claim)Not stated on the pages read.No numeric vendor claim found; marketed as low-latency. Text-LLM cascades are described by Microsoft as having slightly higher latency than native realtime models.No numeric vendor claim captured.
Key limits
  • Concurrent requests: 5 (On-Demand), 10 (Creator), 50 (Builder), 150 (Developer), 500 (Growth), custom (Enterprise)
  • 100,000 tokens per minute per resource by default (increase on request)
  • Max session duration and concurrency are not stated in Voice Live docs
  • Phrase list under 500 words/phrases
  • SIP is not supported directly; use Azure Communication Services for telephony
  • Single WebSocket session up to 120 minutes
  • qwen3.8-omni-flash-realtime: up to 196,608 input tokens; retains 100 audio turns, 50 video turns, 600 s of audio and 240 s of video
  • qwen3.5-omni-flash-realtime retains 80 audio turns, 480 s audio, 120 s video
  • Base64 images under 256 KB; max 1080p
High-severity warnings
  • LLM cost is extra
  • You pay for speech twice in cascaded mode
  • Context replay drives the bill
  • Model ids churn fast
ComplianceNot stated on the pages read.Azure AI services enterprise terms; content filtering always on; data zone and regional model variants for residency. HIPAA BAA via Microsoft for in-scope Azure services (confirm Voice Live scope with Microsoft).Data retention, residency guarantees and certifications for Model Studio realtime were not captured; review Alibaba Cloud International terms before sending regulated data.
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min---
Audio in $/1M tok-$32$0.93
Audio out $/1M tok-$64$1.87
Free tierYesNoYes
Free credit $---
Native S2SNoYesYes
ToolsYesYesYes
Image in-YesYes
Own LLM-YesNo
Voices-600-
CloningYesYesYes
Latency ms---
Languages-14636
Context tokens--196,608
Max session min--120
Concurrency5--
WebRTCYesYesYes
WebSocketYesYesYes
Phone / SIPNoNoNo
HIPAA-Yes-
SOC 2---
EU data-YesNo
Open weightsNoNoNo
High warnings112