Voice-to-voice
One model listens and talks back, or a hosted agent API that bundles speech-to-text, an LLM and a voice for you.
As of 2026-10-10 Anthropic does not offer a realtime or speech-to-speech API. Current Claude models (Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 5.5) take text and image input and return text only; Claude voice mode exists only in Anthropic's own apps.
Neither Meta's Llama API nor Mistral offers a native single-model speech-to-speech realtime API as of 2026-10-10. Mistral documents a voice-agent pipeline instead: Voxtral Realtime (streaming STT, Feb 2026) + an LLM + Voxtral TTS (Mar 2026), with open weights available.
Columns 9/31
Showing 20 of 20. Click any column header to sort.
Qwen-Omni RealtimeAlibaba Cloud |
GA | $0.0018 | $0.015 | - | $0.93 | $1.87 | Yes | - | Yes | Yes | Yes | No | - | Yes | - | 36 | 196,608 | 120 | - | Yes | Yes | No | - | - | No | No | 2 | 13 | Alibaba Cloud | medium | 2026-10-10 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Inworld Realtime APIInworld AI |
GA | $0.01 | $0.03 | - | - | - | Yes | - | No | Yes | - | - | - | Yes | - | - | - | - | 5 | Yes | Yes | No | - | - | - | No | 1 | 5 | Inworld AI | medium | 2026-10-10 |
Azure Speech Voice Live APIMicrosoft |
GA | $0.017 | $0.76 | - | $32 | $64 | No | - | Yes | Yes | Yes | Yes | 600 | Yes | - | 146 | - | - | - | Yes | Yes | No | Yes | - | Yes | No | 1 | 8 | Microsoft | medium | 2026-10-10 |
Amazon Nova SonicAmazon Web Services |
GA | $0.022 | $0.12 | - | $3 | $12 | No | - | Yes | Yes | - | No | 16 | - | - | 7 | 1,000,000 | 8 | 20 | No | No | No | Yes | Yes | Yes | No | 2 | 8 | Amazon Web Services | medium | 2026-10-10 |
Gemini LiveGoogle |
GA | $0.022 | $0.12 | - | $3 | $12 | Yes | - | Yes | Yes | Yes | No | 30 | - | - | 99 | 131,072 | 10 | - | No | Yes | No | No | - | No | No | 3 | 8 | high | 2026-10-10 | |
Gemini Live on VertexGoogle Cloud |
GA | $0.022 | $0.12 | - | $3 | $12 | No | - | Yes | Yes | Yes | No | 30 | - | - | 24 | 128,000 | 10 | 1,000 | No | Yes | No | Yes | - | Yes | No | 2 | 6 | Google Cloud | medium | 2026-10-10 |
GLM-RealtimeZhipu AI (bigmodel.cn) |
GA | $0.025 | $0.30 | $0.025 | - | - | - | - | Yes | Yes | Yes | No | 7 | - | - | - | 8,000 | - | 5 | No | Yes | No | - | - | No | No | 1 | 6 | Zhipu AI (bigmodel.cn) | medium | 2026-10-10 |
Doubao RealtimeByteDance Volcengine (Doubao Speech) |
GA | $0.03 | $0.07 | - | $11.27 | $42.25 | Yes | - | Yes | Yes | - | No | - | Yes | - | 2 | - | - | - | No | Yes | No | - | - | No | No | 2 | 7 | ByteDance Volcengine (Doubao Speech) | medium | 2026-10-10 |
Deepgram Voice AgentDeepgram |
GA | $0.041 | $0.163 | $0.075 | - | - | Yes | $200 | No | Yes | - | Yes | - | - | - | - | - | - | 45 | No | Yes | No | - | - | - | No | 2 | 7 | Deepgram | high | 2026-10-10 |
OpenAI GPT-Live APIOpenAI |
GA | $0.05 | $0.05 | $0.05 | - | - | No | - | Yes | Yes | No | Yes | 12 | - | - | - | 128,000 | - | 50 | Yes | Yes | Yes | - | - | Yes | No | 2 | 8 | OpenAI | medium | 2026-10-10 |
UltravoxUltravox (formerly Fixie.ai) |
GA | $0.05 | $0.055 | $0.05 | - | - | Yes | - | No | Yes | - | - | - | Yes | - | - | - | - | 5 | Yes | Yes | Yes | - | - | - | Yes | 1 | 6 | Ultravox (formerly Fixie.ai) | high | 2026-10-10 |
Cartesia AgentsCartesia |
GA | $0.06 | $0.09 | $0.06 | - | - | - | - | No | Yes | - | - | - | - | - | - | - | - | - | No | Yes | Yes | - | - | - | No | 1 | 5 | Cartesia | high | 2026-10-10 |
AssemblyAI Voice Agent APIAssemblyAI |
GA | $0.075 | $0.075 | $0.075 | - | - | Yes | $50 | No | Yes | - | - | - | - | - | - | - | - | - | No | Yes | No | - | - | - | No | 0 | 5 | AssemblyAI | medium | 2026-10-10 |
ElevenLabs AgentsElevenLabs |
GA | $0.08 | $0.13 | $0.08 | - | - | Yes | - | No | Yes | - | Yes | 5,000 | Yes | - | 31 | - | - | 6 | - | Yes | Yes | Yes | - | Yes | No | 2 | 7 | ElevenLabs | high | 2026-10-10 |
| GA | $0.08 | $0.08 | $0.08 | - | - | No | - | Yes | Yes | - | No | - | Yes | - | - | - | - | - | No | Yes | No | Yes | Yes | Yes | No | 1 | 6 | xAI | medium | 2026-10-10 | |
Azure OpenAI RealtimeMicrosoft |
GA | $0.096 | $0.84 | - | $32 | $64 | No | - | Yes | Yes | Yes | No | 10 | No | - | - | 128,000 | 60 | - | Yes | Yes | Yes | Yes | - | Yes | No | 2 | 7 | Microsoft | medium | 2026-10-10 |
OpenAI RealtimeOpenAI |
GA | $0.096 | $0.76 | - | $32 | $64 | No | - | Yes | Yes | Yes | No | 10 | No | - | - | 128,000 | 60 | - | Yes | Yes | Yes | Yes | Yes | Yes | No | 3 | 8 | OpenAI | high | 2026-10-10 |
PhonicPhonic |
GA | $0.15 | $0.15 | $0.15 | - | - | - | - | Yes | Yes | - | No | - | - | 500 | 51 | - | - | - | No | Yes | Yes | Yes | Yes | - | No | 0 | 5 | Phonic | medium | 2026-10-10 |
StepAudio RealtimeStepFun |
GA | - | - | - | $1.5 | $10 | - | - | Yes | - | - | No | - | Yes | - | - | - | - | - | No | Yes | No | - | - | - | No | 1 | 5 | StepFun | medium | 2026-10-10 |
Hume EVIHume AI |
Deprecated | $0.04 | $0.07 | $0.07 | - | - | Yes | - | Yes | - | - | Yes | - | Yes | - | 11 | - | 30 | - | No | Yes | No | - | - | - | No | 2 | 6 | Hume AI | medium | 2026-10-10 |
A dash means the vendor does not say. Latency is the vendor's own claim, not our measurement. Per-minute figures are estimates from list prices; each API page explains the basis, because vendors bill by tokens, characters, connection time or flat minutes.
Category warnings
These apply to most APIs in this category.
History is re-billed every turn
Every token-billed speech-to-speech API here (OpenAI, Azure, Gemini Live, Vertex Live, Alibaba Qwen-Omni, very likely Nova Sonic) sends the whole conversation, including earlier audio, back through the model on each turn. Cost per minute climbs with call length. In our 10 minute model call, OpenAI gpt-realtime-2.1 goes from about $0.10/min (single turn) to about $0.76/min with no cache hits. Use caching where offered, context truncation/compression, and short system prompts.
Headline per-minute numbers are not comparable
OpenAI audio is 10 tokens/s in and 20 tokens/s out, Gemini Live is 25 tokens/s both ways, Qwen-Omni 3.8 is 7 in and 12.5 out, Nova Sonic does not publish a rate. Per-minute vendors (xAI $0.08/min, GPT-Live $0.05/min plus backend) bill wall-clock session time. Model your own traffic: user talk ratio, turns per minute, call length, tools.
Every session has a hard clock
OpenAI Realtime and Azure: 60 min per session. Gemini Live: ~10 min per connection, 15 min audio-only session without compression (2 min with video). Nova Sonic: 8 min per connection. Alibaba Qwen-Omni: 120 min. Build reconnect plus context carry-over from day one, and do it at a turn boundary.
Never put the long-lived key in the browser
OpenAI and Azure use /realtime/client_secrets ephemeral keys, Gemini uses auth_tokens ephemeral tokens (v1beta), xAI uses /v1/realtime/client_secrets with a sec-websocket-protocol prefix. Nova Sonic has no browser token flow: proxy through your backend. Ephemeral tokens can still be abused while valid, so lock the session config server side and keep expiry short.
Hume EVI and TTS shut down on 2026-11-13
Hume's docs and changelog (notice dated 2026-10-02) say access ends November 13, 2026 at 12:01 a.m. EST and account data is deleted afterwards. Anyone still on EVI must migrate now.
Most 'voice agent' APIs are cascaded, not native speech-to-speech
ElevenLabs, Deepgram, AssemblyAI, Cartesia, Inworld and Kyutai Unmute chain STT, an LLM and TTS. Native speech-to-speech here: Alibaba Qwen-Omni Realtime, Volcengine Doubao, Zhipu GLM-Realtime, StepFun, Phonic, and open models like Moshi, PersonaPlex and MiniCPM-o. Ultravox is audio-native on input but speaks through TTS.
Per-minute prices often exclude the LLM
ElevenLabs ($0.08/min), Cartesia ($0.06/min) and Inworld bill the LLM separately. Deepgram Standard/Advanced, AssemblyAI ($0.075/min) and Ultravox ($0.05/min) include it. Cartesia's free-LLM promotion ended on 2026-10-01.
Echo and barge-in are your problem
When audio plays through speakers the mic hears the agent and it interrupts itself. Only Azure Voice Live offers server-side echo cancellation. Elsewhere rely on browser getUserMedia echoCancellation, WebRTC, or headsets, and on interruption truncate the server-side transcript to what the user actually heard (OpenAI: conversation.item.truncate).
Fast model churn and forced migrations
In 2026 OpenAI removed the Realtime beta interface (May 12), shut down gpt-4o realtime previews (May 7) and will shut gpt-realtime and gpt-realtime-mini on Jan 20, 2027. Google moved from 2.5 native audio to 3.1 Flash Live preview to 3.8 Live within six months. Pin model versions, watch deprecation pages, and budget a migration every few months.
Data residency is uneven
OpenAI EU residency for /v1/realtime needs approved abuse-monitoring controls; Azure has Global vs Data Zone deployments (Data Zone costs 10 percent more); Nova Sonic is in-region only in 4 regions; Gemini Developer API has no region choice, Vertex does. Check before promising customers where audio is processed.
China-region providers need China accounts and pay in CNY
Volcengine Doubao, Zhipu bigmodel.cn, StepFun (.com) and Alibaba Beijing are mainland-China services: expect real-name verification, Chinese-language consoles and data processed in China. Only Alibaba (Singapore) has a clearly documented international realtime region in this segment.
Connection time is what you pay for
Deepgram, ElevenLabs, Ultravox and AssemblyAI meter session or connection minutes. Idle sockets, unanswered outbound calls and default 1-hour max durations (Ultravox) can quietly burn money. Always set max duration and inactivity timeouts.
No native speech-to-speech from Anthropic, Meta or Mistral
As of 2026-10-10 Anthropic, Meta (Llama API) and Mistral do not offer a single-model realtime speech-to-speech API. With those LLMs you build a cascade (streaming STT + LLM + streaming TTS) or use a platform that does it for you.
Prices quoted in CNY are converted roughly
USD estimates for Chinese providers use about 7.1 CNY per USD and are approximations, not vendor figures.