Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Doubao Realtime Voice Model (end-to-end realtime speech)
ByteDance Volcengine (Doubao Speech)
GLM-Realtime
Zhipu AI (bigmodel.cn)
Gemini Live API (Gemini Developer API / Google AI Studio)
Google
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minute$0.03 - 0.07$0.025 - 0.30$0.022 - 0.12
How that was worked outOwn estimate at about 7.1 CNY per USD: 30 s user + 30 s agent speech is about 0.015 CNY input + 0.225 CNY output audio plus cached context, about 0.25-0.35 CNY/min; an agent speaking the full minute is about 0.45 CNY of output audio.Conversion at about 7.1 CNY per USD: flash audio about $0.025, air audio about $0.042, air video about $0.30 per minute.Low = gemini-3.8-live, 1 min user audio (1,500 tokens x $3/1M) + 1 min model audio (1,500 tokens x $12/1M), single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Live models do not support context caching, so there is no cached discount. Video input or thinking tokens add more.
Pricing modelper-tokenper-minuteper-token
Free tierFree quota exists and can offset cached, uncached and output tokens, but the size was not found in the docs read.Not stated on the GLM-Realtime page.Yes: free tier for all Live models with lower rate limits; free-tier content may be used to improve Google products.
Connects byWebSocketWebSocketWebSocket
Audio inPCM 16 kHz mono int16 little-endian (Opus also accepted and converted server-side); send 20 ms / 640-byte packets at real-time pacewav or pcm (pcm16 = 16 kHz, pcm24 = 24 kHz), mono 16-bitRaw 16-bit PCM little-endian, natively 16 kHz (other rates resampled if the MIME type says so, e.g. audio/pcm;rate=16000). Images/video as JPEG or PNG frames, max 1 fps.
Audio outOgg Opus by default; PCM 24 kHz mono (32-bit float or s16le) on requestPCM 24 kHz mono 16-bitRaw 16-bit PCM little-endian at 24 kHz (always).
LanguagesChinese and English (vendor says other languages are not guaranteed, especially for cloned voices).Multilingual with automatic language detection (no list published); replies in the user's language.Live guide lists 99 languages (overview page says 70); native audio models pick the language automatically and do not accept a language code.
Latency (vendor claim)Vendor describes it as low latency; no millisecond figure on the API page.Not published.No numeric vendor claim found.
Key limits
  • Default QPM 60 (StartSession / session.create per minute per AppID) and TPM 100,000; raise via sales
  • Server releases the connection after 10 minutes with no interaction (error 45000003)
  • Uplink audio must keep real-time pace; send input_audio_mute.commit when the mic is muted or the session times out
  • Close with session.close and wait for the reply, otherwise error 55000001 ContextCanceled
  • Context 8K for audio calls (about 20 turns per docs) and 32K for video calls
  • Conversation memory up to about 2 minutes
  • max_response_output_tokens up to 1024
  • Client VAD mode: max 30 s per upload; send at most 50 messages per second (100 ms frames recommended)
  • Connection lifetime around 10 minutes (GoAway with timeLeft is sent before close)
  • Without compression: audio-only sessions 15 minutes, audio+video 2 minutes
  • Context window 128k tokens for native audio models (3.8 Live lists 131,072 input)
  • Session resumption tokens valid 2 hours after the last session ends (Developer API)
High-severity warnings
  • China-only account and KYC
  • Two incompatible protocols
  • Very short memory
  • 10 minute connections, 15 minute sessions
  • Whole context re-billed every turn, no caching
  • Free tier trains on your data
ComplianceNot stated for this API in the docs read.Not stated on the page read.Paid tier data is not used to improve products; free tier is. No BAA or data residency on the Developer API; use Vertex AI for enterprise compliance (CMEK, VPC-SC, regional processing).
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min-$0.025-
Audio in $/1M tok$11.27-$3
Audio out $/1M tok$42.25-$12
Free tierYes-Yes
Free credit $---
Native S2SYesYesYes
ToolsYesYesYes
Image in-YesYes
Own LLMNoNoNo
Voices-730
CloningYes--
Latency ms---
Languages2-99
Context tokens-8,000131,072
Max session min--10
Concurrency-5-
WebRTCNoNoNo
WebSocketYesYesYes
Phone / SIPNoNoNo
HIPAA--No
SOC 2---
EU dataNoNoNo
Open weightsNoNoNo
High warnings213