Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Alibaba Cloud Model Studio Qwen-Omni-Realtime
Alibaba Cloud
Inworld Realtime API
Inworld AI
Azure Speech Voice Live API
Microsoft
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minute$0.0018 - 0.015$0.01 - 0.03$0.017 - 0.76
How that was worked outLow = qwen3.8-omni-flash-realtime, 1 min user audio (420 tokens x $0.93/1M) + 1 min model audio (750 tokens x $1.87/1M), single turn, excluding the matching output text tokens (small). High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Alibaba documents that each turn re-bills retained history. qwen3.5-omni-plus-realtime: $0.053 low, $0.27 high.Own estimate: STT about $0.0025/min plus TTS-2 about $0.0125-0.025 per minute of agent speech (1,000 chars/min), before LLM cost.Low = Lite tier phi4-mm-realtime, 1 min user audio (750 tokens x $4/1M) + 1 min model audio (1,200 tokens x $12/1M), single turn. Pro tier with gpt-realtime-2.1 native audio matches Azure OpenAI Global: $0.096 single turn. High = Pro gpt-realtime-2.1, High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Standard tier gpt-realtime-2.1-mini: $0.033 low, $0.26 high. Avatar minutes and custom voice are extra.
Pricing modelper-tokenper-characterper-token
Free tier1M free tokens per model in the Singapore region for 90 days.On-Demand plan free to start with up to 70 TTS minutes and up to 400 STT minutes.No dedicated free tier found for Voice Live (Azure Speech F0 is limited to one concurrent request).
Connects byWebSocket, WebRTCWebSocket, WebRTCWebSocket, WebRTC
Audio inPCM 16 kHz (16-bit mono LE); multichannel 2 or 4 channel spatial input supported (double tokens); JPG images about 1 fps for video.PCM16 mono 24 kHz default; G.711 mu-law/A-law 8 kHz; float32PCM16 mono at 24 kHz (default) or 16 kHz via input_audio_sampling_rate; Live-Reference AEC mode takes interleaved stereo PCM16 (mic + playback reference). Fixed for the session.
Audio outPCM 24 kHz.Same optionsPCM16 audio from the native model or Azure TTS; word timestamps and visemes available with Azure voices; avatar video over WebRTC (H.264).
LanguagesSpeech recognition 113 languages and dialects; speech generation 36 languages and dialects.Depends on STT/TTS models chosen; not listed on the pages read.146 input locales and 151 output locales per the FAQ (Azure speech); azure_semantic_vad_multilingual covers English, Spanish, French, Italian, German, Japanese, Portuguese, Chinese, Korean, Hindi.
Latency (vendor claim)No numeric vendor claim captured.Not stated on the pages read.No numeric vendor claim found; marketed as low-latency. Text-LLM cascades are described by Microsoft as having slightly higher latency than native realtime models.
Key limits
  • Single WebSocket session up to 120 minutes
  • qwen3.8-omni-flash-realtime: up to 196,608 input tokens; retains 100 audio turns, 50 video turns, 600 s of audio and 240 s of video
  • qwen3.5-omni-flash-realtime retains 80 audio turns, 480 s audio, 120 s video
  • Base64 images under 256 KB; max 1080p
  • Concurrent requests: 5 (On-Demand), 10 (Creator), 50 (Builder), 150 (Developer), 500 (Growth), custom (Enterprise)
  • 100,000 tokens per minute per resource by default (increase on request)
  • Max session duration and concurrency are not stated in Voice Live docs
  • Phrase list under 500 words/phrases
  • SIP is not supported directly; use Azure Communication Services for telephony
High-severity warnings
  • Context replay drives the bill
  • Model ids churn fast
  • LLM cost is extra
  • You pay for speech twice in cascaded mode
ComplianceData retention, residency guarantees and certifications for Model Studio realtime were not captured; review Alibaba Cloud International terms before sending regulated data.Not stated on the pages read.Azure AI services enterprise terms; content filtering always on; data zone and regional model variants for residency. HIPAA BAA via Microsoft for in-scope Azure services (confirm Voice Live scope with Microsoft).
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min---
Audio in $/1M tok$0.93-$32
Audio out $/1M tok$1.87-$64
Free tierYesYesNo
Free credit $---
Native S2SYesNoYes
ToolsYesYesYes
Image inYes-Yes
Own LLMNo-Yes
Voices--600
CloningYesYesYes
Latency ms---
Languages36-146
Context tokens196,608--
Max session min120--
Concurrency-5-
WebRTCYesYesYes
WebSocketYesYesYes
Phone / SIPNoNoNo
HIPAA--Yes
SOC 2---
EU dataNo-Yes
Open weightsNoNoNo
High warnings211