Read this before you build
The traps that cost teams money, a rewrite or a launch date, grouped by category. Each API page has its own list on top of these.
Voice-to-voice
Browse 20 APIsHistory is re-billed every turn
Every token-billed speech-to-speech API here (OpenAI, Azure, Gemini Live, Vertex Live, Alibaba Qwen-Omni, very likely Nova Sonic) sends the whole conversation, including earlier audio, back through the model on each turn. Cost per minute climbs with call length. In our 10 minute model call, OpenAI gpt-realtime-2.1 goes from about $0.10/min (single turn) to about $0.76/min with no cache hits. Use caching where offered, context truncation/compression, and short system prompts.
Headline per-minute numbers are not comparable
OpenAI audio is 10 tokens/s in and 20 tokens/s out, Gemini Live is 25 tokens/s both ways, Qwen-Omni 3.8 is 7 in and 12.5 out, Nova Sonic does not publish a rate. Per-minute vendors (xAI $0.08/min, GPT-Live $0.05/min plus backend) bill wall-clock session time. Model your own traffic: user talk ratio, turns per minute, call length, tools.
Every session has a hard clock
OpenAI Realtime and Azure: 60 min per session. Gemini Live: ~10 min per connection, 15 min audio-only session without compression (2 min with video). Nova Sonic: 8 min per connection. Alibaba Qwen-Omni: 120 min. Build reconnect plus context carry-over from day one, and do it at a turn boundary.
Never put the long-lived key in the browser
OpenAI and Azure use /realtime/client_secrets ephemeral keys, Gemini uses auth_tokens ephemeral tokens (v1beta), xAI uses /v1/realtime/client_secrets with a sec-websocket-protocol prefix. Nova Sonic has no browser token flow: proxy through your backend. Ephemeral tokens can still be abused while valid, so lock the session config server side and keep expiry short.
Hume EVI and TTS shut down on 2026-11-13
Hume's docs and changelog (notice dated 2026-10-02) say access ends November 13, 2026 at 12:01 a.m. EST and account data is deleted afterwards. Anyone still on EVI must migrate now.
Most 'voice agent' APIs are cascaded, not native speech-to-speech
ElevenLabs, Deepgram, AssemblyAI, Cartesia, Inworld and Kyutai Unmute chain STT, an LLM and TTS. Native speech-to-speech here: Alibaba Qwen-Omni Realtime, Volcengine Doubao, Zhipu GLM-Realtime, StepFun, Phonic, and open models like Moshi, PersonaPlex and MiniCPM-o. Ultravox is audio-native on input but speaks through TTS.
Per-minute prices often exclude the LLM
ElevenLabs ($0.08/min), Cartesia ($0.06/min) and Inworld bill the LLM separately. Deepgram Standard/Advanced, AssemblyAI ($0.075/min) and Ultravox ($0.05/min) include it. Cartesia's free-LLM promotion ended on 2026-10-01.
Echo and barge-in are your problem
When audio plays through speakers the mic hears the agent and it interrupts itself. Only Azure Voice Live offers server-side echo cancellation. Elsewhere rely on browser getUserMedia echoCancellation, WebRTC, or headsets, and on interruption truncate the server-side transcript to what the user actually heard (OpenAI: conversation.item.truncate).
Fast model churn and forced migrations
In 2026 OpenAI removed the Realtime beta interface (May 12), shut down gpt-4o realtime previews (May 7) and will shut gpt-realtime and gpt-realtime-mini on Jan 20, 2027. Google moved from 2.5 native audio to 3.1 Flash Live preview to 3.8 Live within six months. Pin model versions, watch deprecation pages, and budget a migration every few months.
Data residency is uneven
OpenAI EU residency for /v1/realtime needs approved abuse-monitoring controls; Azure has Global vs Data Zone deployments (Data Zone costs 10 percent more); Nova Sonic is in-region only in 4 regions; Gemini Developer API has no region choice, Vertex does. Check before promising customers where audio is processed.
China-region providers need China accounts and pay in CNY
Volcengine Doubao, Zhipu bigmodel.cn, StepFun (.com) and Alibaba Beijing are mainland-China services: expect real-name verification, Chinese-language consoles and data processed in China. Only Alibaba (Singapore) has a clearly documented international realtime region in this segment.
Connection time is what you pay for
Deepgram, ElevenLabs, Ultravox and AssemblyAI meter session or connection minutes. Idle sockets, unanswered outbound calls and default 1-hour max durations (Ultravox) can quietly burn money. Always set max duration and inactivity timeouts.
No native speech-to-speech from Anthropic, Meta or Mistral
As of 2026-10-10 Anthropic, Meta (Llama API) and Mistral do not offer a single-model realtime speech-to-speech API. With those LLMs you build a cascade (streaming STT + LLM + streaming TTS) or use a platform that does it for you.
Prices quoted in CNY are converted roughly
USD estimates for Chinese providers use about 7.1 CNY per USD and are approximations, not vendor figures.
Speech-to-text, live
Browse 19 APIsBilling basis differs: socket time vs audio vs tokens
AssemblyAI bills the time the WebSocket is open; Rev AI bills the larger of stream time and audio time with a 15 s minimum; Cartesia bills audio seconds including silence; Soniox and older OpenAI models bill tokens. An idle open connection can cost the same as speech, so close sessions as soon as a call ends.
Idle timeouts and keepalives
Deepgram closes after 10 s with no audio or KeepAlive; Speechmatics after 3 minutes without audio or ping; AssemblyAI has an inactivity_timeout and a 3-hour cap. Mute buttons, hold music gaps and slow LLM turns are the usual triggers. Send keepalives or silence and handle reconnects.
Maximum stream length
Google V2 streaming is about 5 minutes per stream; AssemblyAI, Gladia and Rev AI cap at 3 hours; Soniox 300 minutes (fixed); Speechmatics 48 hours. Long meetings need overlapping reconnects and transcript stitching.
Partial vs final results
Partials (interim, non-final tokens) are rewritten as more audio arrives. Only commit finals to storage, LLM context or analytics; render partials as provisional UI. Some APIs send deltas (Cartesia, OpenAI) that must be concatenated exactly as received.
Endpointing is a latency vs cut-off trade-off
Short silence thresholds make voice agents snappy but split sentences at natural pauses; long thresholds add hundreds of ms per turn. Turn-aware models (Deepgram Flux, AssemblyAI turn detection, Speechmatics Agent STT, Smallest pulse-2, Soniox semantic endpointing) help, but tune per use case and language.
Streaming diarization is often missing or extra
Google Chirp 3 (batch only), Mistral Voxtral Realtime, ElevenLabs Scribe Realtime and Gladia live do not document live speaker labels; Deepgram and AssemblyAI charge extra on streaming. For phone calls, send each party as its own channel instead.
8 kHz telephony audio
Send phone audio at its native 8 kHz mu-law/PCM when the API supports it (Deepgram mulaw, ElevenLabs ulaw_8000, Speechmatics mulaw, Cartesia pcm_mulaw, Google MULAW) instead of upsampling; accuracy on narrowband audio is lower than vendor benchmarks, so test with real calls.
Never ship long-lived keys to browsers
Use short-lived credentials: Deepgram JWT (30 s TTL), AssemblyAI temporary tokens, Speechmatics JWT, ElevenLabs single-use tokens, Mistral rt_ tokens (~900 s), Cartesia access tokens, OpenAI ephemeral client secrets, Gladia per-session URLs. Some APIs (Rev AI, Sarvam) only offer key-in-URL or subprotocol auth, so relay through your server.
Real-time pacing and chunk size
Most APIs expect audio at about real-time speed in 20-200 ms chunks; AssemblyAI rejects chunks outside 50-1000 ms or faster-than-real-time sends, Deepgram Flux wants 80 ms, AWS recommends 50-200 ms. Streaming a file needs a sleep between chunks.
Rapid model churn in 2026
Soniox v4 to v5 (auto-routed), AssemblyAI moved to universal-3-6-pro and dropped u3-rt-pro IDs, OpenAI added gpt-realtime-whisper then gpt-live-transcribe and gpt-transcribe, Fireworks deprecated audio. Pin model IDs, watch changelogs and re-test before upgrading.
Promotional and unlabeled prices
Deepgram shows streaming rates as limited-time promotional, Azure lists MAI-Transcribe-2 under a promotion to 2026-12-31, and several vendors (Rev AI, Sarvam) do not separate streaming from batch prices. Confirm the streaming line item in writing.
Latency claims are not comparable
Vendors measure differently (first partial, end-of-turn, final). Benchmark time-to-final on your own audio, network and region instead of comparing marketing numbers.
Text-to-speech, streaming
Browse 26 APIsVendors are disappearing in 2025-2026
PlayHT's API went offline in July 2025, LMNT's site now says it has shut down, Hume ends its TTS and EVI APIs on 2026-11-13, Groq retired PlayAI TTS, and Rime retired Arcana. Put TTS behind your own interface, keep the source audio for any cloned voice, and keep a second provider tested.
Units differ: characters, credits, UTF-8 bytes, tokens
ElevenLabs and Cartesia bill credits per character, Fish Audio per UTF-8 byte (non-Latin scripts cost ~3x), OpenAI, Gemini-TTS and Soniox per audio token (25 tokens/s for Gemini), Inworld counts UTF-16 code units. Convert everything to cost per minute of your real content before comparing.
Promotional and preview prices expire
ElevenLabs v4/v4 Turbo launch prices end Oct 12 2026 (roughly 3.6x higher afterwards), Gemini 3.8 Flash TTS prices double on Jan 1 2027, Fish Audio's free s2.1-pro-free is reported to end Nov 30 2026, Unreal Speech Basic is $4.99 only for 6 months. Budget on list prices.
Whitespace, markup and SSML are billed
Google counts spaces, newlines and SSML tags (except <mark>); Cartesia charges 1 credit per break tag. Strip markdown, emoji and extra whitespace from LLM output before sending it to TTS.
Vendor latency claims are not your latency
Most headline numbers are model or server-side TTFB excluding network (e.g. Inworld 20 ms P90 server-side, Murf 55 ms model). Independent medians including network are typically 100-450 ms (Vapi Humanness Index, Coval). Measure time-to-first-audio from your own region and also total turn latency.
Chunk LLM output at sentence boundaries
Even with WebSocket input streaming, most engines buffer until punctuation or a character threshold (ElevenLabs chunk_length_schedule starts at 120 chars, MiniMax waits for sentence-final punctuation, ElevenLabs Text to Dialogue waits for ~40 chars and 8 words). Send text at clause/sentence ends and send an explicit flush at end of turn; do not flush every token (Deepgram allows only 20 flushes/min).
HTTP-only providers need your own sentence splitter
OpenAI, Speechify, Unreal Speech, Groq (200-char cap) and Resemble's socket accept a whole text per request. For agents, split replies into sentences and pipeline requests; watch per-request caps and rate limits.
Concurrency limits bite before volume does
Entry plans allow very few simultaneous generations: Smallest.ai 1 per account, Murf 2 outside US-East, Cartesia Free 2 / Pro 3, Fish Audio 5 until $100 spent, Inworld On-Demand 5. Size plans by peak simultaneous speakers.
Voice cloning: consent and disclosure
Clone only voices you have written permission to use. OpenAI requires a recorded consent statement for custom voices and disclosure that the voice is AI-generated; open models (Chatterbox, Dia, Nari Labs) prohibit impersonation. Disclosure duties for synthetic audio also exist in some jurisdictions (for example the EU AI Act transparency rules).
Ask for telephony formats natively
For phone calls request mu-law or A-law 8 kHz from the provider (Cartesia, Deepgram, ElevenLabs, Inworld, Rime, Google, MiniMax support it) instead of resampling. Avoid WAV headers inside streams: they click (Deepgram recommends container=none; Inworld LINEAR16 repeats a header per chunk).
Browsers cannot set WebSocket auth headers
Rime, Deepgram, Inworld, Fish and others expect a header. Proxy through your backend or use the vendor's short-lived tokens (Cartesia access_token, ElevenLabs single_use_token). Never ship a long-lived key to the client; Murf puts the key in the URL, so keep that server-side.
Logging and retention defaults
ElevenLabs enable_logging defaults to true; Fish's free model may retain requests for training. Use zero-retention options (Inworld, Smallest enterprise, Rime retention controls) for sensitive text.
Platforms and telephony
Browse 23 APIsHeadline per-minute price is almost never the all-in price
Most platforms quote a platform or orchestration fee, then pass through speech-to-text, LLM and text-to-speech costs, then add telephony on top. Example: Vapi's $0.05/min is hosting only; its own estimator puts 1,000 minutes at $82-$129 before Twilio/Telnyx minutes. Retell's $0.055/min voice infra excludes the LLM, TTS upgrades and $0.015/min telephony. Budget = platform fee + STT + LLM + TTS + carrier minutes + phone number rental + add-ons (knowledge base, denoising, PII redaction, branded caller ID).
FCC 2024 ruling: AI voices count as 'artificial voice' under the TCPA
The FCC Declaratory Ruling of 8 Feb 2024 confirmed that calls using AI-generated or cloned voices are 'artificial or prerecorded voice' calls under the TCPA. Outbound AI calls to consumers need prior express consent (prior express written consent for telemarketing), must identify the calling entity, and telemarketing calls must offer an opt-out. There is no carve-out for AI that behaves like a live agent. Violations expose you to FCC fines, state AG suits and private TCPA class actions. Inbound calls where the consumer called you are a different, lower-risk situation. Not legal advice; check current rules and state laws (several states add their own AI-disclosure or telemarketing rules).
Call recording and AI disclosure consent
Several US states (for example California, Florida, Pennsylvania, Washington, Illinois) require all-party consent to record, and most voice-agent platforms record and transcribe by default. In the EU/UK, recordings and transcripts are personal data under GDPR. Play a short disclosure at the start of the call ('this call is with an AI assistant and may be recorded') and turn recording off where you do not need it.
HIPAA needs a BAA with every hop, not just the platform
A platform BAA does not cover the STT, LLM, TTS or carrier vendors it calls on your behalf unless they are inside the BAA scope. Some platforms charge heavily for it (Vapi lists HIPAA as a $2,000/month add-on; Thoughtly offers HIPAA/BAA only on Enterprise). Confirm the exact list of covered sub-processors before sending PHI.
Billing rounding, minimums and silence
Telnyx estimates exclude per-call rounding to 60-second increments; Agora rounds SIP/PSTN agent calls up to the next full minute; Bland charges a $0.015 minimum per outbound attempt and per failed call on its telephony; Retell bills per second but keeps billing during silence and hold. Short, high-volume outbound campaigns can cost several times the headline rate. Set max call duration, silence timeouts and voicemail detection so you do not pay agent minutes to talk to an answering machine.
Concurrency caps on entry tiers
Entry plans cap simultaneous calls: Vapi 4 on the free package (10 on Core, $10/line/month extra), Retell 20 included ($8 per extra concurrent call per month), Bland 10 on Start, LiveKit Cloud 5 agent sessions on Build, Thoughtly 10 on Flex. Campaign dialers hit these immediately; calls over the cap are queued or rejected.
Outbound numbers get spam-labelled
Carriers score outbound numbers with STIR/SHAKEN attestation plus call analytics. New numbers that place many short calls quickly show as 'Spam Likely' and answer rates collapse. Register numbers with the carrier analytics registries, keep volume per number reasonable, use numbers your platform can sign with full (A) attestation, and consider paid branded calling (Twilio lists $0.12/call, Retell $0.10/outbound call).
Phone audio is 8 kHz narrowband
PSTN legs are usually G.711 mu-law at 8 kHz (Twilio Media Streams, many SIP trunks). STT accuracy drops versus 16 kHz web audio, TTS is downsampled, and you must resample correctly (sending 16 kHz or 24 kHz PCM into an 8 kHz mu-law stream produces chipmunk or garbled audio). Pick telephony-tuned STT models and test with real phone calls, not browser demos. Plivo and some others offer 16 kHz linear PCM streams if your trunk supports it.
WebRTC needs TURN, and corporate networks block UDP
Browser and mobile agents fail silently behind strict NAT or corporate firewalls without a TURN relay (ideally TURN over TLS on port 443). Managed SFUs (LiveKit Cloud, Daily, Cloudflare, Agora, Stream) include this; self-hosted setups must run and pay for TURN egress themselves.
Echo and noise cause self-interruption
On speakerphones and laptop speakers the agent hears its own TTS and barges in on itself, or background noise triggers false turns. Keep browser echoCancellation/noiseSuppression on, use server-side noise cancellation (Krisp in Pipecat Cloud, ai-coustics/BVC in LiveKit, Retell 'Advanced Denoising' at +$0.005/min), and tune interruption sensitivity.
Vendor churn is real in this category
PlayAI was acquired by Meta in July 2025 and play.ai no longer resolves (checked 2026-10-10). layercode.com and docs.layercode.com now redirect to toyo.ai. The FTC sued Air AI in August 2025 over claims about its AI sales agents. Synthflow removed public self-serve pricing (Enterprise from $30,000/year). Keep prompts, tools and call logs portable, and prefer frameworks (Pipecat, LiveKit Agents) or platforms with export APIs if lock-in worries you.
Latency figures are vendor claims
Every vendor publishes a best-case latency number. Real response time depends on the region of the telephony edge, the agent worker, and each model endpoint. Co-locate them and measure p50/p95 end-of-speech to first-audio on real calls before choosing.
Avatars and live video
Browse 22 APIsCustom avatars need the real person's consent
Vendors that clone a real face require a recorded consent statement from that person (Tavus requires a word-for-word consent statement in the training video for personal replicas, and a matching consent video if you use footage of someone else). Never build a custom avatar from footage of someone who has not agreed, and keep the consent records.
EU AI Act Article 50 disclosure applies since 2 August 2026
Secondary legal commentary reports that Article 50 transparency duties (tell users they are talking to an AI, label deepfakes) took effect on 2 August 2026 and were not delayed by the Digital Omnibus; only machine-readable marking for systems already on the market got a grace period to 2 December 2026. Show a visible 'AI avatar' notice in any realtime avatar or video-to-video product used in the EU. Not legal advice; check the official text.
The session meter runs while nobody is talking
Avatar vendors bill wall-clock session time, not speaking time. Anam says silence, a muted mic or an idle avatar does not pause billing; bitHuman bills active session time whether talking or idle; Tavus bills a 30 second minimum per conversation. Always set a max session length and an idle or participant-left timeout, and end sessions server-side.
Concurrency caps are low on cheap plans
Entry plans usually allow 1 to 3 simultaneous sessions (Tavus Free/Starter 1, Anam Free/Starter 1, LemonSlice Starter 3, bitHuman Creator 3). Over the cap you get HTTP 429 or a burst surcharge (LemonSlice bills calls above the limit at 2x). Size the plan by peak concurrent users, not by total minutes.
Avatar price is only one layer of the bill
Most avatar APIs only render the face. Speech-to-text, the LLM, text-to-speech and the WebRTC room (LiveKit, Daily, Agora) are billed separately unless you buy the vendor's managed 'full' mode. Budget the whole pipeline per minute.
Realtime video-to-video is far pricier than avatars
Generative realtime video is billed per second of GPU output: Decart Lucy 2.5 is $0.02/sec at 720p ($1.20 per minute, $72 per hour), double in fast mode. Leaving a stream open in a background tab can burn money quickly; stop streams on tab blur and enforce a hard cap.
Products in this space are being retired fast
Hedra's Realtime Avatar was sunset on 15 April 2026 (LiveKit marks the plugin as no longer working), HeyGen's Interactive/Streaming Avatar API was replaced by LiveAvatar with a stated sunset of 31 March 2026, and Soul Machines went into receivership in February 2026. Keep the avatar layer behind an adapter so you can swap vendors.
Video ramps up and needs bandwidth
WebRTC video takes time to stabilise (Roboflow says quality and FPS can take up to a minute to ramp at 1080p30), and corporate firewalls can block UDP. Test on mobile data and behind a strict firewall, and provide TURN.
No realtime API from Luma, Twelve Labs or Google Genie
As of this research Luma's API is asynchronous job-based, Twelve Labs streams text output from analysis of stored video (no live-feed ingest found), and Google's Project Genie is a consumer research prototype for AI Ultra subscribers with no developer API. Runway's realtime offer is Runway Characters (avatars); GWM Worlds 2 is contact-form preview only per secondary reporting.
Open models
Browse 30 APIsOpen-weight licences are not all permissive
Apache-2.0/MIT: Qwen3-Omni, MiniCPM-o 4.5, Step-Audio 2 mini, Kimi-Audio, Fun-Audio-Chat, Sesame CSM, Unmute code. CC-BY-4.0 (attribution): Moshi and Kyutai STT/TTS. Custom: NVIDIA Open Model License (PersonaPlex), GLM-4 licence (GLM-4-Voice), LFM Open License with a $10M revenue cap (Liquid LFM2.5-Audio).
Open models rarely ship a production realtime server
Only Moshi/PersonaPlex (moshi.server), Unmute (docker compose) and MiniCPM-o (llama.cpp-omni full duplex) come close. Qwen3-Omni's vLLM path did not output speech at release. Budget engineering time for VAD, streaming, barge-in, auth and scaling.
Free tiers and open weights are not always commercial
Gradium Free is explicitly non-commercial; F5-TTS weights are CC-BY-NC, XTTS-v2 uses the non-commercial Coqui licence (no one left to sell a commercial one), Mistral Voxtral TTS weights and Fish Audio S2 Pro weights are non-commercial. Apache-2.0/MIT options: Kokoro, Chatterbox, Qwen3-TTS, VoxCPM2, Dia, Sesame CSM, NeuTTS Air.
Shutdown watch
| API | Status |
|---|---|
| Hume EVI (Empathic Voice Interface) | Deprecated |
| Fireworks AI streaming ASR (deprecated) | Deprecated |
| Hume Octave TTS | Deprecated |
| LMNT | Shut down |
| PlayHT (Play.ai) | Shut down |
| Coqui XTTS-v2 | Deprecated |
| Layercode | Deprecated |
| PlayAI Agents | Deprecated |
| Hedra Realtime Avatar | Deprecated |
| Soul Machines | Deprecated |
Most high-severity warnings
Live APIs with the longest list of serious catches. Not a quality ranking: big platforms simply have more surface.