All the realtime AI APIs. Prices, setup, pitfalls.
140 voice, speech, avatar and video APIs side by side. See what a minute really costs, copy a working quick start, and read the warnings before you build.
Voice-to-voice, upper cost per minute
green under $0.05 / amber under $0.20 / red above. Scale tops out at $0.80.
Upper end of each range. For token-billed models that is a 10 minute call where the history is billed again every turn; flat per-minute APIs stay flat.
Channels
Six categories. Each has a sortable table, side-by-side compare and a full page per API.
Voice-to-voice
One model listens and talks back, or a hosted agent API that bundles speech-to-text, an LLM...
19Speech-to-text, live
Streaming transcription with partial results and turn detection, for captions, notes and voice...
26Text-to-speech, streaming
Voices that start speaking in a few hundred milliseconds, ideally while the LLM is still...
23Platforms and telephony
Frameworks, media transport, phone lines and hosted agent builders that tie the pieces together.
22Avatars and live video
Talking faces, realtime video models and live vision input.
30Open models
Weights you can run on your own GPU. Hardware needs and licence terms are listed for each one.
Alarm panel
What breaks budgets and launches. See all 689 warnings.
History is re-billed every turn
Every token-billed speech-to-speech API here (OpenAI, Azure, Gemini Live, Vertex Live, Alibaba Qwen-Omni, very likely Nova Sonic) sends the whole conversation, including earlier audio, back through the model on each turn. Cost...
Never put the long-lived key in the browser
OpenAI and Azure use /realtime/client_secrets ephemeral keys, Gemini uses auth_tokens ephemeral tokens (v1beta), xAI uses /v1/realtime/client_secrets with a sec-websocket-protocol prefix. Nova Sonic has no browser token flow:...
Headline per-minute price is almost never the all-in price
Most platforms quote a platform or orchestration fee, then pass through speech-to-text, LLM and text-to-speech costs, then add telephony on top. Example: Vapi's $0.05/min is hosting only; its own estimator puts 1,000 minutes at...
FCC 2024 ruling: AI voices count as 'artificial voice' under the TCPA
The FCC Declaratory Ruling of 8 Feb 2024 confirmed that calls using AI-generated or cloned voices are 'artificial or prerecorded voice' calls under the TCPA. Outbound AI calls to consumers need prior express consent (prior...
Billing basis differs: socket time vs audio vs tokens
AssemblyAI bills the time the WebSocket is open; Rev AI bills the larger of stream time and audio time with a 15 s minimum; Cartesia bills audio seconds including silence; Soniox and older OpenAI models bill tokens. An idle open...
The session meter runs while nobody is talking
Avatar vendors bill wall-clock session time, not speaking time. Anam says silence, a muted mic or an idle avatar does not pause billing; bitHuman bills active session time whether talking or idle; Tavus bills a 30 second minimum...
Lowest list price per category
A starting point, not a recommendation. Open each page for the catches.
Shutdown watch
APIs that are retired, deprecated or about to go dark.
| API | Status |
|---|---|
| Hume EVI (Empathic Voice Interface) Hume AI | Deprecated |
| Fireworks AI streaming ASR (deprecated) Fireworks AI | Deprecated |
| Hume Octave TTS Hume AI | Deprecated |
| LMNT LMNT | Shut down |
| PlayHT (Play.ai) PlayHT (team acqui-hired by Meta) | Shut down |
| Coqui XTTS-v2 Coqui (defunct); maintained fork by Idiap | Deprecated |
| Layercode Layercode, Inc. | Deprecated |
| PlayAI Agents PlayAI (acquired by Meta) | Deprecated |
| Hedra Realtime Avatar Hedra | Deprecated |
| Soul Machines Soul Machines (in receivership) | Deprecated |
Questions people ask
What is a realtime AI API?
An API that streams audio or video both ways while a conversation is happening, so a model can listen, think and answer in well under a second. It usually runs over a WebSocket, WebRTC or a phone line (SIP) instead of one request and one response.
How much does a voice AI agent cost per minute?
It depends on the stack. By list price, a voice-to-voice model ranges from about $0.0018 a minute (Alibaba Cloud Model Studio Qwen-Omni-Realtime) to roughly $0.76 a minute for a long OpenAI gpt-realtime-2.1 call. Hosted agent platforms usually land between $0.07 and $0.31 a minute once speech-to-text, the LLM, the voice and the phone line are added up.
Why does a long voice call cost more per minute than a short one?
Token-billed voice models send the whole conversation back through the model on every turn, including earlier audio. The longer the call, the more each new turn costs, unless you use caching, truncation or context compression.
Should I use one voice-to-voice model or a speech-to-text, LLM and text-to-speech pipeline?
Voice-to-voice models sound more natural and react faster, but you get fewer voices and less control. A pipeline lets you pick any LLM and voice, log clean transcripts and swap parts, at the cost of more moving pieces and usually a bit more latency.
How current are these prices?
Every provider page shows the date it was last checked against the vendor's own pricing and docs, and links to its sources. The full set was checked on Oct 10, 2026. Entries older than 60 days are flagged as possibly out of date.