Voice-to-voice

One model listens and talks back, or a hosted agent API that bundles speech-to-text, an LLM and a voice for you.

Anthropic Claude (no realtime voice API)

As of 2026-10-10 Anthropic does not offer a realtime or speech-to-speech API. Current Claude models (Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 5.5) take text and image input and return text only; Claude voice mode exists only in Anthropic's own apps.

Meta Llama API and Mistral (no single-model realtime S2S)

Neither Meta's Llama API nor Mistral offers a native single-model speech-to-speech realtime API as of 2026-10-10. Mistral documents a voice-agent pipeline instead: Voxtral Realtime (streaming STT, Feb 2026) + an LLM + Voxtral TTS (Mar 2026), with open weights available.

Columns 9/31
Basics
Cost
Features
Performance
Limits
Connects by
Compliance
Openness
Warnings
Compare 0
Must have:

Showing 20 of 20. Click any column header to sort.

Qwen-Omni RealtimeAlibaba Cloud
GA$0.0018$0.015YesYesYes-362
Inworld Realtime APIInworld AI
GA$0.01$0.03YesNoYes--1
Azure Speech Voice Live APIMicrosoft
GA$0.017$0.76NoYesYes-1461
Amazon Nova SonicAmazon Web Services
GA$0.022$0.12NoYesYes-72
Gemini LiveGoogle
GA$0.022$0.12YesYesYes-993
Gemini Live on VertexGoogle Cloud
GA$0.022$0.12NoYesYes-242
GLM-RealtimeZhipu AI (bigmodel.cn)
GA$0.025$0.30-YesYes--1
Doubao RealtimeByteDance Volcengine (Doubao Speech)
GA$0.03$0.07YesYesYes-22
Deepgram Voice AgentDeepgram
GA$0.041$0.163YesNoYes--2
OpenAI GPT-Live APIOpenAI
GA$0.05$0.05NoYesYes--2
UltravoxUltravox (formerly Fixie.ai)
GA$0.05$0.055YesNoYes--1
Cartesia AgentsCartesia
GA$0.06$0.09-NoYes--1
AssemblyAI Voice Agent APIAssemblyAI
GA$0.075$0.075YesNoYes--0
ElevenLabs AgentsElevenLabs
GA$0.08$0.13YesNoYes-312
Grok Voice AgentxAI
GA$0.08$0.08NoYesYes--1
Azure OpenAI RealtimeMicrosoft
GA$0.096$0.84NoYesYes--2
OpenAI RealtimeOpenAI
GA$0.096$0.76NoYesYes--3
PhonicPhonic
GA$0.15$0.15-YesYes500510
StepAudio RealtimeStepFun
GA---Yes---1
Hume EVIHume AI
Deprecated$0.04$0.07YesYes--112

A dash means the vendor does not say. Latency is the vendor's own claim, not our measurement. Per-minute figures are estimates from list prices; each API page explains the basis, because vendors bill by tokens, characters, connection time or flat minutes.

Category warnings

These apply to most APIs in this category.

History is re-billed every turn

Every token-billed speech-to-speech API here (OpenAI, Azure, Gemini Live, Vertex Live, Alibaba Qwen-Omni, very likely Nova Sonic) sends the whole conversation, including earlier audio, back through the model on each turn. Cost per minute climbs with call length. In our 10 minute model call, OpenAI gpt-realtime-2.1 goes from about $0.10/min (single turn) to about $0.76/min with no cache hits. Use caching where offered, context truncation/compression, and short system prompts.

Headline per-minute numbers are not comparable

OpenAI audio is 10 tokens/s in and 20 tokens/s out, Gemini Live is 25 tokens/s both ways, Qwen-Omni 3.8 is 7 in and 12.5 out, Nova Sonic does not publish a rate. Per-minute vendors (xAI $0.08/min, GPT-Live $0.05/min plus backend) bill wall-clock session time. Model your own traffic: user talk ratio, turns per minute, call length, tools.

Every session has a hard clock

OpenAI Realtime and Azure: 60 min per session. Gemini Live: ~10 min per connection, 15 min audio-only session without compression (2 min with video). Nova Sonic: 8 min per connection. Alibaba Qwen-Omni: 120 min. Build reconnect plus context carry-over from day one, and do it at a turn boundary.

Never put the long-lived key in the browser

OpenAI and Azure use /realtime/client_secrets ephemeral keys, Gemini uses auth_tokens ephemeral tokens (v1beta), xAI uses /v1/realtime/client_secrets with a sec-websocket-protocol prefix. Nova Sonic has no browser token flow: proxy through your backend. Ephemeral tokens can still be abused while valid, so lock the session config server side and keep expiry short.

Hume EVI and TTS shut down on 2026-11-13

Hume's docs and changelog (notice dated 2026-10-02) say access ends November 13, 2026 at 12:01 a.m. EST and account data is deleted afterwards. Anyone still on EVI must migrate now.

Most 'voice agent' APIs are cascaded, not native speech-to-speech

ElevenLabs, Deepgram, AssemblyAI, Cartesia, Inworld and Kyutai Unmute chain STT, an LLM and TTS. Native speech-to-speech here: Alibaba Qwen-Omni Realtime, Volcengine Doubao, Zhipu GLM-Realtime, StepFun, Phonic, and open models like Moshi, PersonaPlex and MiniCPM-o. Ultravox is audio-native on input but speaks through TTS.

Per-minute prices often exclude the LLM

ElevenLabs ($0.08/min), Cartesia ($0.06/min) and Inworld bill the LLM separately. Deepgram Standard/Advanced, AssemblyAI ($0.075/min) and Ultravox ($0.05/min) include it. Cartesia's free-LLM promotion ended on 2026-10-01.

Echo and barge-in are your problem

When audio plays through speakers the mic hears the agent and it interrupts itself. Only Azure Voice Live offers server-side echo cancellation. Elsewhere rely on browser getUserMedia echoCancellation, WebRTC, or headsets, and on interruption truncate the server-side transcript to what the user actually heard (OpenAI: conversation.item.truncate).

Fast model churn and forced migrations

In 2026 OpenAI removed the Realtime beta interface (May 12), shut down gpt-4o realtime previews (May 7) and will shut gpt-realtime and gpt-realtime-mini on Jan 20, 2027. Google moved from 2.5 native audio to 3.1 Flash Live preview to 3.8 Live within six months. Pin model versions, watch deprecation pages, and budget a migration every few months.

Data residency is uneven

OpenAI EU residency for /v1/realtime needs approved abuse-monitoring controls; Azure has Global vs Data Zone deployments (Data Zone costs 10 percent more); Nova Sonic is in-region only in 4 regions; Gemini Developer API has no region choice, Vertex does. Check before promising customers where audio is processed.

China-region providers need China accounts and pay in CNY

Volcengine Doubao, Zhipu bigmodel.cn, StepFun (.com) and Alibaba Beijing are mainland-China services: expect real-name verification, Chinese-language consoles and data processed in China. Only Alibaba (Singapore) has a clearly documented international realtime region in this segment.

Connection time is what you pay for

Deepgram, ElevenLabs, Ultravox and AssemblyAI meter session or connection minutes. Idle sockets, unanswered outbound calls and default 1-hour max durations (Ultravox) can quietly burn money. Always set max duration and inactivity timeouts.

No native speech-to-speech from Anthropic, Meta or Mistral

As of 2026-10-10 Anthropic, Meta (Llama API) and Mistral do not offer a single-model realtime speech-to-speech API. With those LLMs you build a cascade (streaming STT + LLM + streaming TTS) or use a platform that does it for you.

Prices quoted in CNY are converted roughly

USD estimates for Chinese providers use about 7.1 CNY per USD and are approximations, not vendor figures.