Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Doubao Realtime Voice Model (end-to-end realtime speech) ByteDance Volcengine (Doubao Speech) | GLM-Realtime Zhipu AI (bigmodel.cn) | Gemini Live API (Gemini Developer API / Google AI Studio) Google | |
|---|---|---|---|
| Category | Voice-to-voice | Voice-to-voice | Voice-to-voice |
| Status | GA | GA | GA |
| Est. per minute | $0.03 - 0.07 | $0.025 - 0.30 | $0.022 - 0.12 |
| How that was worked out | Own estimate at about 7.1 CNY per USD: 30 s user + 30 s agent speech is about 0.015 CNY input + 0.225 CNY output audio plus cached context, about 0.25-0.35 CNY/min; an agent speaking the full minute is about 0.45 CNY of output audio. | Conversion at about 7.1 CNY per USD: flash audio about $0.025, air audio about $0.042, air video about $0.30 per minute. | Low = gemini-3.8-live, 1 min user audio (1,500 tokens x $3/1M) + 1 min model audio (1,500 tokens x $12/1M), single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Live models do not support context caching, so there is no cached discount. Video input or thinking tokens add more. |
| Pricing model | per-token | per-minute | per-token |
| Free tier | Free quota exists and can offset cached, uncached and output tokens, but the size was not found in the docs read. | Not stated on the GLM-Realtime page. | Yes: free tier for all Live models with lower rate limits; free-tier content may be used to improve Google products. |
| Connects by | WebSocket | WebSocket | WebSocket |
| Audio in | PCM 16 kHz mono int16 little-endian (Opus also accepted and converted server-side); send 20 ms / 640-byte packets at real-time pace | wav or pcm (pcm16 = 16 kHz, pcm24 = 24 kHz), mono 16-bit | Raw 16-bit PCM little-endian, natively 16 kHz (other rates resampled if the MIME type says so, e.g. audio/pcm;rate=16000). Images/video as JPEG or PNG frames, max 1 fps. |
| Audio out | Ogg Opus by default; PCM 24 kHz mono (32-bit float or s16le) on request | PCM 24 kHz mono 16-bit | Raw 16-bit PCM little-endian at 24 kHz (always). |
| Languages | Chinese and English (vendor says other languages are not guaranteed, especially for cloned voices). | Multilingual with automatic language detection (no list published); replies in the user's language. | Live guide lists 99 languages (overview page says 70); native audio models pick the language automatically and do not accept a language code. |
| Latency (vendor claim) | Vendor describes it as low latency; no millisecond figure on the API page. | Not published. | No numeric vendor claim found. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Not stated for this API in the docs read. | Not stated on the page read. | Paid tier data is not used to improve products; free tier is. No BAA or data residency on the Developer API; use Vertex AI for enterprise compliance (CMEK, VPC-SC, regional processing). |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Flat $/min | - | $0.025 | - |
| Audio in $/1M tok | $11.27 | - | $3 |
| Audio out $/1M tok | $42.25 | - | $12 |
| Free tier | Yes | - | Yes |
| Free credit $ | - | - | - |
| Native S2S | Yes | Yes | Yes |
| Tools | Yes | Yes | Yes |
| Image in | - | Yes | Yes |
| Own LLM | No | No | No |
| Voices | - | 7 | 30 |
| Cloning | Yes | - | - |
| Latency ms | - | - | - |
| Languages | 2 | - | 99 |
| Context tokens | - | 8,000 | 131,072 |
| Max session min | - | - | 10 |
| Concurrency | - | 5 | - |
| WebRTC | No | No | No |
| WebSocket | Yes | Yes | Yes |
| Phone / SIP | No | No | No |
| HIPAA | - | - | No |
| SOC 2 | - | - | - |
| EU data | No | No | No |
| Open weights | No | No | No |
| High warnings | 2 | 1 | 3 |