{"name":"Realtime API Index","url":"https://realtime.anri.ai","verified_at":"2026-10-10","licence":"Free to use with a link back to realtime.anri.ai","categories":[{"id":"voice","name":"Voice-to-voice","short":"Voice-to-voice","blurb":"One model listens and talks back, or a hosted agent API that bundles speech-to-text, an LLM and a voice for you.","warnings":[{"severity":"high","title":"History is re-billed every turn","detail":"Every token-billed speech-to-speech API here (OpenAI, Azure, Gemini Live, Vertex Live, Alibaba Qwen-Omni, very likely Nova Sonic) sends the whole conversation, including earlier audio, back through the model on each turn. Cost per minute climbs with call length. In our 10 minute model call, OpenAI gpt-realtime-2.1 goes from about $0.10/min (single turn) to about $0.76/min with no cache hits. Use caching where offered, context truncation/compression, and short system prompts."},{"severity":"high","title":"Headline per-minute numbers are not comparable","detail":"OpenAI audio is 10 tokens/s in and 20 tokens/s out, Gemini Live is 25 tokens/s both ways, Qwen-Omni 3.8 is 7 in and 12.5 out, Nova Sonic does not publish a rate. Per-minute vendors (xAI $0.08/min, GPT-Live $0.05/min plus backend) bill wall-clock session time. Model your own traffic: user talk ratio, turns per minute, call length, tools."},{"severity":"high","title":"Every session has a hard clock","detail":"OpenAI Realtime and Azure: 60 min per session. Gemini Live: ~10 min per connection, 15 min audio-only session without compression (2 min with video). Nova Sonic: 8 min per connection. Alibaba Qwen-Omni: 120 min. Build reconnect plus context carry-over from day one, and do it at a turn boundary."},{"severity":"high","title":"Never put the long-lived key in the browser","detail":"OpenAI and Azure use /realtime/client_secrets ephemeral keys, Gemini uses auth_tokens ephemeral tokens (v1beta), xAI uses /v1/realtime/client_secrets with a sec-websocket-protocol prefix. Nova Sonic has no browser token flow: proxy through your backend. Ephemeral tokens can still be abused while valid, so lock the session config server side and keep expiry short."},{"severity":"high","title":"Hume EVI and TTS shut down on 2026-11-13","detail":"Hume's docs and changelog (notice dated 2026-10-02) say access ends November 13, 2026 at 12:01 a.m. EST and account data is deleted afterwards. Anyone still on EVI must migrate now."},{"severity":"high","title":"Most 'voice agent' APIs are cascaded, not native speech-to-speech","detail":"ElevenLabs, Deepgram, AssemblyAI, Cartesia, Inworld and Kyutai Unmute chain STT, an LLM and TTS. Native speech-to-speech here: Alibaba Qwen-Omni Realtime, Volcengine Doubao, Zhipu GLM-Realtime, StepFun, Phonic, and open models like Moshi, PersonaPlex and MiniCPM-o. Ultravox is audio-native on input but speaks through TTS."},{"severity":"high","title":"Per-minute prices often exclude the LLM","detail":"ElevenLabs ($0.08/min), Cartesia ($0.06/min) and Inworld bill the LLM separately. Deepgram Standard/Advanced, AssemblyAI ($0.075/min) and Ultravox ($0.05/min) include it. Cartesia's free-LLM promotion ended on 2026-10-01."},{"severity":"medium","title":"Echo and barge-in are your problem","detail":"When audio plays through speakers the mic hears the agent and it interrupts itself. Only Azure Voice Live offers server-side echo cancellation. Elsewhere rely on browser getUserMedia echoCancellation, WebRTC, or headsets, and on interruption truncate the server-side transcript to what the user actually heard (OpenAI: conversation.item.truncate)."},{"severity":"medium","title":"Fast model churn and forced migrations","detail":"In 2026 OpenAI removed the Realtime beta interface (May 12), shut down gpt-4o realtime previews (May 7) and will shut gpt-realtime and gpt-realtime-mini on Jan 20, 2027. Google moved from 2.5 native audio to 3.1 Flash Live preview to 3.8 Live within six months. Pin model versions, watch deprecation pages, and budget a migration every few months."},{"severity":"medium","title":"Data residency is uneven","detail":"OpenAI EU residency for /v1/realtime needs approved abuse-monitoring controls; Azure has Global vs Data Zone deployments (Data Zone costs 10 percent more); Nova Sonic is in-region only in 4 regions; Gemini Developer API has no region choice, Vertex does. Check before promising customers where audio is processed."},{"severity":"medium","title":"China-region providers need China accounts and pay in CNY","detail":"Volcengine Doubao, Zhipu bigmodel.cn, StepFun (.com) and Alibaba Beijing are mainland-China services: expect real-name verification, Chinese-language consoles and data processed in China. Only Alibaba (Singapore) has a clearly documented international realtime region in this segment."},{"severity":"medium","title":"Connection time is what you pay for","detail":"Deepgram, ElevenLabs, Ultravox and AssemblyAI meter session or connection minutes. Idle sockets, unanswered outbound calls and default 1-hour max durations (Ultravox) can quietly burn money. Always set max duration and inactivity timeouts."},{"severity":"low","title":"No native speech-to-speech from Anthropic, Meta or Mistral","detail":"As of 2026-10-10 Anthropic, Meta (Llama API) and Mistral do not offer a single-model realtime speech-to-speech API. With those LLMs you build a cascade (streaming STT + LLM + streaming TTS) or use a platform that does it for you."},{"severity":"low","title":"Prices quoted in CNY are converted roughly","detail":"USD estimates for Chinese providers use about 7.1 CNY per USD and are approximations, not vendor figures."}]},{"id":"stt","name":"Speech-to-text, live","short":"Speech-to-text","blurb":"Streaming transcription with partial results and turn detection, for captions, notes and voice agents.","warnings":[{"severity":"high","title":"Billing basis differs: socket time vs audio vs tokens","detail":"AssemblyAI bills the time the WebSocket is open; Rev AI bills the larger of stream time and audio time with a 15 s minimum; Cartesia bills audio seconds including silence; Soniox and older OpenAI models bill tokens. An idle open connection can cost the same as speech, so close sessions as soon as a call ends."},{"severity":"high","title":"Idle timeouts and keepalives","detail":"Deepgram closes after 10 s with no audio or KeepAlive; Speechmatics after 3 minutes without audio or ping; AssemblyAI has an inactivity_timeout and a 3-hour cap. Mute buttons, hold music gaps and slow LLM turns are the usual triggers. Send keepalives or silence and handle reconnects."},{"severity":"high","title":"Maximum stream length","detail":"Google V2 streaming is about 5 minutes per stream; AssemblyAI, Gladia and Rev AI cap at 3 hours; Soniox 300 minutes (fixed); Speechmatics 48 hours. Long meetings need overlapping reconnects and transcript stitching."},{"severity":"medium","title":"Partial vs final results","detail":"Partials (interim, non-final tokens) are rewritten as more audio arrives. Only commit finals to storage, LLM context or analytics; render partials as provisional UI. Some APIs send deltas (Cartesia, OpenAI) that must be concatenated exactly as received."},{"severity":"medium","title":"Endpointing is a latency vs cut-off trade-off","detail":"Short silence thresholds make voice agents snappy but split sentences at natural pauses; long thresholds add hundreds of ms per turn. Turn-aware models (Deepgram Flux, AssemblyAI turn detection, Speechmatics Agent STT, Smallest pulse-2, Soniox semantic endpointing) help, but tune per use case and language."},{"severity":"medium","title":"Streaming diarization is often missing or extra","detail":"Google Chirp 3 (batch only), Mistral Voxtral Realtime, ElevenLabs Scribe Realtime and Gladia live do not document live speaker labels; Deepgram and AssemblyAI charge extra on streaming. For phone calls, send each party as its own channel instead."},{"severity":"medium","title":"8 kHz telephony audio","detail":"Send phone audio at its native 8 kHz mu-law/PCM when the API supports it (Deepgram mulaw, ElevenLabs ulaw_8000, Speechmatics mulaw, Cartesia pcm_mulaw, Google MULAW) instead of upsampling; accuracy on narrowband audio is lower than vendor benchmarks, so test with real calls."},{"severity":"medium","title":"Never ship long-lived keys to browsers","detail":"Use short-lived credentials: Deepgram JWT (30 s TTL), AssemblyAI temporary tokens, Speechmatics JWT, ElevenLabs single-use tokens, Mistral rt_ tokens (~900 s), Cartesia access tokens, OpenAI ephemeral client secrets, Gladia per-session URLs. Some APIs (Rev AI, Sarvam) only offer key-in-URL or subprotocol auth, so relay through your server."},{"severity":"medium","title":"Real-time pacing and chunk size","detail":"Most APIs expect audio at about real-time speed in 20-200 ms chunks; AssemblyAI rejects chunks outside 50-1000 ms or faster-than-real-time sends, Deepgram Flux wants 80 ms, AWS recommends 50-200 ms. Streaming a file needs a sleep between chunks."},{"severity":"medium","title":"Rapid model churn in 2026","detail":"Soniox v4 to v5 (auto-routed), AssemblyAI moved to universal-3-6-pro and dropped u3-rt-pro IDs, OpenAI added gpt-realtime-whisper then gpt-live-transcribe and gpt-transcribe, Fireworks deprecated audio. Pin model IDs, watch changelogs and re-test before upgrading."},{"severity":"low","title":"Promotional and unlabeled prices","detail":"Deepgram shows streaming rates as limited-time promotional, Azure lists MAI-Transcribe-2 under a promotion to 2026-12-31, and several vendors (Rev AI, Sarvam) do not separate streaming from batch prices. Confirm the streaming line item in writing."},{"severity":"low","title":"Latency claims are not comparable","detail":"Vendors measure differently (first partial, end-of-turn, final). Benchmark time-to-final on your own audio, network and region instead of comparing marketing numbers."}]},{"id":"tts","name":"Text-to-speech, streaming","short":"Text-to-speech","blurb":"Voices that start speaking in a few hundred milliseconds, ideally while the LLM is still writing.","warnings":[{"severity":"high","title":"Vendors are disappearing in 2025-2026","detail":"PlayHT's API went offline in July 2025, LMNT's site now says it has shut down, Hume ends its TTS and EVI APIs on 2026-11-13, Groq retired PlayAI TTS, and Rime retired Arcana. Put TTS behind your own interface, keep the source audio for any cloned voice, and keep a second provider tested."},{"severity":"high","title":"Units differ: characters, credits, UTF-8 bytes, tokens","detail":"ElevenLabs and Cartesia bill credits per character, Fish Audio per UTF-8 byte (non-Latin scripts cost ~3x), OpenAI, Gemini-TTS and Soniox per audio token (25 tokens/s for Gemini), Inworld counts UTF-16 code units. Convert everything to cost per minute of your real content before comparing."},{"severity":"high","title":"Promotional and preview prices expire","detail":"ElevenLabs v4/v4 Turbo launch prices end Oct 12 2026 (roughly 3.6x higher afterwards), Gemini 3.8 Flash TTS prices double on Jan 1 2027, Fish Audio's free s2.1-pro-free is reported to end Nov 30 2026, Unreal Speech Basic is $4.99 only for 6 months. Budget on list prices."},{"severity":"medium","title":"Whitespace, markup and SSML are billed","detail":"Google counts spaces, newlines and SSML tags (except <mark>); Cartesia charges 1 credit per break tag. Strip markdown, emoji and extra whitespace from LLM output before sending it to TTS."},{"severity":"medium","title":"Vendor latency claims are not your latency","detail":"Most headline numbers are model or server-side TTFB excluding network (e.g. Inworld 20 ms P90 server-side, Murf 55 ms model). Independent medians including network are typically 100-450 ms (Vapi Humanness Index, Coval). Measure time-to-first-audio from your own region and also total turn latency."},{"severity":"medium","title":"Chunk LLM output at sentence boundaries","detail":"Even with WebSocket input streaming, most engines buffer until punctuation or a character threshold (ElevenLabs chunk_length_schedule starts at 120 chars, MiniMax waits for sentence-final punctuation, ElevenLabs Text to Dialogue waits for ~40 chars and 8 words). Send text at clause/sentence ends and send an explicit flush at end of turn; do not flush every token (Deepgram allows only 20 flushes/min)."},{"severity":"medium","title":"HTTP-only providers need your own sentence splitter","detail":"OpenAI, Speechify, Unreal Speech, Groq (200-char cap) and Resemble's socket accept a whole text per request. For agents, split replies into sentences and pipeline requests; watch per-request caps and rate limits."},{"severity":"medium","title":"Concurrency limits bite before volume does","detail":"Entry plans allow very few simultaneous generations: Smallest.ai 1 per account, Murf 2 outside US-East, Cartesia Free 2 / Pro 3, Fish Audio 5 until $100 spent, Inworld On-Demand 5. Size plans by peak simultaneous speakers."},{"severity":"medium","title":"Voice cloning: consent and disclosure","detail":"Clone only voices you have written permission to use. OpenAI requires a recorded consent statement for custom voices and disclosure that the voice is AI-generated; open models (Chatterbox, Dia, Nari Labs) prohibit impersonation. Disclosure duties for synthetic audio also exist in some jurisdictions (for example the EU AI Act transparency rules)."},{"severity":"low","title":"Ask for telephony formats natively","detail":"For phone calls request mu-law or A-law 8 kHz from the provider (Cartesia, Deepgram, ElevenLabs, Inworld, Rime, Google, MiniMax support it) instead of resampling. Avoid WAV headers inside streams: they click (Deepgram recommends container=none; Inworld LINEAR16 repeats a header per chunk)."},{"severity":"low","title":"Browsers cannot set WebSocket auth headers","detail":"Rime, Deepgram, Inworld, Fish and others expect a header. Proxy through your backend or use the vendor's short-lived tokens (Cartesia access_token, ElevenLabs single_use_token). Never ship a long-lived key to the client; Murf puts the key in the URL, so keep that server-side."},{"severity":"low","title":"Logging and retention defaults","detail":"ElevenLabs enable_logging defaults to true; Fish's free model may retain requests for training. Use zero-retention options (Inworld, Smallest enterprise, Rime retention controls) for sensitive text."}]},{"id":"platforms","name":"Platforms and telephony","short":"Platforms","blurb":"Frameworks, media transport, phone lines and hosted agent builders that tie the pieces together.","warnings":[{"severity":"high","title":"Headline per-minute price is almost never the all-in price","detail":"Most platforms quote a platform or orchestration fee, then pass through speech-to-text, LLM and text-to-speech costs, then add telephony on top. Example: Vapi's $0.05/min is hosting only; its own estimator puts 1,000 minutes at $82-$129 before Twilio/Telnyx minutes. Retell's $0.055/min voice infra excludes the LLM, TTS upgrades and $0.015/min telephony. Budget = platform fee + STT + LLM + TTS + carrier minutes + phone number rental + add-ons (knowledge base, denoising, PII redaction, branded caller ID)."},{"severity":"high","title":"FCC 2024 ruling: AI voices count as 'artificial voice' under the TCPA","detail":"The FCC Declaratory Ruling of 8 Feb 2024 confirmed that calls using AI-generated or cloned voices are 'artificial or prerecorded voice' calls under the TCPA. Outbound AI calls to consumers need prior express consent (prior express written consent for telemarketing), must identify the calling entity, and telemarketing calls must offer an opt-out. There is no carve-out for AI that behaves like a live agent. Violations expose you to FCC fines, state AG suits and private TCPA class actions. Inbound calls where the consumer called you are a different, lower-risk situation. Not legal advice; check current rules and state laws (several states add their own AI-disclosure or telemarketing rules)."},{"severity":"high","title":"Call recording and AI disclosure consent","detail":"Several US states (for example California, Florida, Pennsylvania, Washington, Illinois) require all-party consent to record, and most voice-agent platforms record and transcribe by default. In the EU/UK, recordings and transcripts are personal data under GDPR. Play a short disclosure at the start of the call ('this call is with an AI assistant and may be recorded') and turn recording off where you do not need it."},{"severity":"high","title":"HIPAA needs a BAA with every hop, not just the platform","detail":"A platform BAA does not cover the STT, LLM, TTS or carrier vendors it calls on your behalf unless they are inside the BAA scope. Some platforms charge heavily for it (Vapi lists HIPAA as a $2,000/month add-on; Thoughtly offers HIPAA/BAA only on Enterprise). Confirm the exact list of covered sub-processors before sending PHI."},{"severity":"medium","title":"Billing rounding, minimums and silence","detail":"Telnyx estimates exclude per-call rounding to 60-second increments; Agora rounds SIP/PSTN agent calls up to the next full minute; Bland charges a $0.015 minimum per outbound attempt and per failed call on its telephony; Retell bills per second but keeps billing during silence and hold. Short, high-volume outbound campaigns can cost several times the headline rate. Set max call duration, silence timeouts and voicemail detection so you do not pay agent minutes to talk to an answering machine."},{"severity":"medium","title":"Concurrency caps on entry tiers","detail":"Entry plans cap simultaneous calls: Vapi 4 on the free package (10 on Core, $10/line/month extra), Retell 20 included ($8 per extra concurrent call per month), Bland 10 on Start, LiveKit Cloud 5 agent sessions on Build, Thoughtly 10 on Flex. Campaign dialers hit these immediately; calls over the cap are queued or rejected."},{"severity":"medium","title":"Outbound numbers get spam-labelled","detail":"Carriers score outbound numbers with STIR/SHAKEN attestation plus call analytics. New numbers that place many short calls quickly show as 'Spam Likely' and answer rates collapse. Register numbers with the carrier analytics registries, keep volume per number reasonable, use numbers your platform can sign with full (A) attestation, and consider paid branded calling (Twilio lists $0.12/call, Retell $0.10/outbound call)."},{"severity":"medium","title":"Phone audio is 8 kHz narrowband","detail":"PSTN legs are usually G.711 mu-law at 8 kHz (Twilio Media Streams, many SIP trunks). STT accuracy drops versus 16 kHz web audio, TTS is downsampled, and you must resample correctly (sending 16 kHz or 24 kHz PCM into an 8 kHz mu-law stream produces chipmunk or garbled audio). Pick telephony-tuned STT models and test with real phone calls, not browser demos. Plivo and some others offer 16 kHz linear PCM streams if your trunk supports it."},{"severity":"medium","title":"WebRTC needs TURN, and corporate networks block UDP","detail":"Browser and mobile agents fail silently behind strict NAT or corporate firewalls without a TURN relay (ideally TURN over TLS on port 443). Managed SFUs (LiveKit Cloud, Daily, Cloudflare, Agora, Stream) include this; self-hosted setups must run and pay for TURN egress themselves."},{"severity":"medium","title":"Echo and noise cause self-interruption","detail":"On speakerphones and laptop speakers the agent hears its own TTS and barges in on itself, or background noise triggers false turns. Keep browser echoCancellation/noiseSuppression on, use server-side noise cancellation (Krisp in Pipecat Cloud, ai-coustics/BVC in LiveKit, Retell 'Advanced Denoising' at +$0.005/min), and tune interruption sensitivity."},{"severity":"low","title":"Vendor churn is real in this category","detail":"PlayAI was acquired by Meta in July 2025 and play.ai no longer resolves (checked 2026-10-10). layercode.com and docs.layercode.com now redirect to toyo.ai. The FTC sued Air AI in August 2025 over claims about its AI sales agents. Synthflow removed public self-serve pricing (Enterprise from $30,000/year). Keep prompts, tools and call logs portable, and prefer frameworks (Pipecat, LiveKit Agents) or platforms with export APIs if lock-in worries you."},{"severity":"low","title":"Latency figures are vendor claims","detail":"Every vendor publishes a best-case latency number. Real response time depends on the region of the telephony edge, the agent worker, and each model endpoint. Co-locate them and measure p50/p95 end-of-speech to first-audio on real calls before choosing."}]},{"id":"video","name":"Avatars and live video","short":"Avatars + video","blurb":"Talking faces, realtime video models and live vision input.","warnings":[{"severity":"high","title":"Custom avatars need the real person's consent","detail":"Vendors that clone a real face require a recorded consent statement from that person (Tavus requires a word-for-word consent statement in the training video for personal replicas, and a matching consent video if you use footage of someone else). Never build a custom avatar from footage of someone who has not agreed, and keep the consent records."},{"severity":"high","title":"EU AI Act Article 50 disclosure applies since 2 August 2026","detail":"Secondary legal commentary reports that Article 50 transparency duties (tell users they are talking to an AI, label deepfakes) took effect on 2 August 2026 and were not delayed by the Digital Omnibus; only machine-readable marking for systems already on the market got a grace period to 2 December 2026. Show a visible 'AI avatar' notice in any realtime avatar or video-to-video product used in the EU. Not legal advice; check the official text."},{"severity":"high","title":"The session meter runs while nobody is talking","detail":"Avatar vendors bill wall-clock session time, not speaking time. Anam says silence, a muted mic or an idle avatar does not pause billing; bitHuman bills active session time whether talking or idle; Tavus bills a 30 second minimum per conversation. Always set a max session length and an idle or participant-left timeout, and end sessions server-side."},{"severity":"medium","title":"Concurrency caps are low on cheap plans","detail":"Entry plans usually allow 1 to 3 simultaneous sessions (Tavus Free/Starter 1, Anam Free/Starter 1, LemonSlice Starter 3, bitHuman Creator 3). Over the cap you get HTTP 429 or a burst surcharge (LemonSlice bills calls above the limit at 2x). Size the plan by peak concurrent users, not by total minutes."},{"severity":"medium","title":"Avatar price is only one layer of the bill","detail":"Most avatar APIs only render the face. Speech-to-text, the LLM, text-to-speech and the WebRTC room (LiveKit, Daily, Agora) are billed separately unless you buy the vendor's managed 'full' mode. Budget the whole pipeline per minute."},{"severity":"medium","title":"Realtime video-to-video is far pricier than avatars","detail":"Generative realtime video is billed per second of GPU output: Decart Lucy 2.5 is $0.02/sec at 720p ($1.20 per minute, $72 per hour), double in fast mode. Leaving a stream open in a background tab can burn money quickly; stop streams on tab blur and enforce a hard cap."},{"severity":"medium","title":"Products in this space are being retired fast","detail":"Hedra's Realtime Avatar was sunset on 15 April 2026 (LiveKit marks the plugin as no longer working), HeyGen's Interactive/Streaming Avatar API was replaced by LiveAvatar with a stated sunset of 31 March 2026, and Soul Machines went into receivership in February 2026. Keep the avatar layer behind an adapter so you can swap vendors."},{"severity":"low","title":"Video ramps up and needs bandwidth","detail":"WebRTC video takes time to stabilise (Roboflow says quality and FPS can take up to a minute to ramp at 1080p30), and corporate firewalls can block UDP. Test on mobile data and behind a strict firewall, and provide TURN."},{"severity":"low","title":"No realtime API from Luma, Twelve Labs or Google Genie","detail":"As of this research Luma's API is asynchronous job-based, Twelve Labs streams text output from analysis of stored video (no live-feed ingest found), and Google's Project Genie is a consumer research prototype for AI Ultra subscribers with no developer API. Runway's realtime offer is Runway Characters (avatars); GWM Worlds 2 is contact-form preview only per secondary reporting."}]},{"id":"open","name":"Open models","short":"Open models","blurb":"Weights you can run on your own GPU. Hardware needs and licence terms are listed for each one.","warnings":[{"severity":"medium","title":"Open-weight licences are not all permissive","detail":"Apache-2.0/MIT: Qwen3-Omni, MiniCPM-o 4.5, Step-Audio 2 mini, Kimi-Audio, Fun-Audio-Chat, Sesame CSM, Unmute code. CC-BY-4.0 (attribution): Moshi and Kyutai STT/TTS. Custom: NVIDIA Open Model License (PersonaPlex), GLM-4 licence (GLM-4-Voice), LFM Open License with a $10M revenue cap (Liquid LFM2.5-Audio)."},{"severity":"medium","title":"Open models rarely ship a production realtime server","detail":"Only Moshi/PersonaPlex (moshi.server), Unmute (docker compose) and MiniCPM-o (llama.cpp-omni full duplex) come close. Qwen3-Omni's vLLM path did not output speech at release. Budget engineering time for VAD, streaming, barge-in, auth and scaling."},{"severity":"medium","title":"Free tiers and open weights are not always commercial","detail":"Gradium Free is explicitly non-commercial; F5-TTS weights are CC-BY-NC, XTTS-v2 uses the non-commercial Coqui licence (no one left to sell a commercial one), Mistral Voxtral TTS weights and Fish Audio S2 Pro weights are non-commercial. Apache-2.0/MIT options: Kokoro, Chatterbox, Qwen3-TTS, VoxCPM2, Dia, Sesame CSM, NeuTTS Air."}]}],"notes":[{"id":"anthropic-no-realtime","name":"Anthropic Claude (no realtime voice API)","vendor":"Anthropic","category":"note","summary":"As of 2026-10-10 Anthropic does not offer a realtime or speech-to-speech API. Current Claude models (Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 5.5) take text and image input and return text only; Claude voice mode exists only in Anthropic's own apps.","status":"GA","models":[],"transports":[],"audio":{"input":"Not supported by the Claude API (the OpenAI-compatibility layer strips audio input).","output":"Not supported."},"languages":"n/a","voices":"n/a","latency":"n/a","features":[],"pricing":{"model":"per-token","items":[],"audio_token_rate":"n/a","est_per_minute_usd":{"low":null,"high":null,"basis":"No audio API. A Claude voice agent is a cascade: streaming STT + Claude (text) + streaming TTS, priced as the sum of those services."},"free_tier":"n/a","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},"limits":[],"regions":"n/a","setup":{"steps":["Use a streaming STT service, send transcripts to the Claude Messages API with streaming, and pipe the text into a streaming TTS service.","Or use a voice platform that supports Claude as the LLM (Anthropic's cookbook shows ElevenLabs STT/TTS with Claude)."],"endpoint":"n/a","auth":"n/a","snippet_lang":"javascript","snippet":""},"warnings":[{"severity":"medium","title":"No native audio","detail":"Do not plan on a Claude speech-to-speech endpoint; voice mode in the Claude apps and Claude Code dictation are not available through the API with an API key."}],"best_for":"Use Claude as the reasoning layer in a cascaded voice pipeline.","open_source":false,"self_hostable":false,"compliance":"n/a","docs":[{"label":"Models overview","url":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"label":"OpenAI SDK compatibility (audio ignored)","url":"https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk"},{"label":"Cookbook: low latency voice assistant with ElevenLabs","url":"https://platform.claude.com/cookbook/third-party-elevenlabs-low-latency-stt-claude-tts"}],"sources":["https://platform.claude.com/docs/en/about-claude/models/overview","https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk","https://code.claude.com/docs/en/voice-dictation"],"confidence":"high","unverified":"An unannounced or private-preview audio API cannot be ruled out.","cat":"voice"},{"id":"meta-mistral-no-native-s2s","name":"Meta Llama API and Mistral (no single-model realtime S2S)","vendor":"Meta, Mistral AI","category":"note","summary":"Neither Meta's Llama API nor Mistral offers a native single-model speech-to-speech realtime API as of 2026-10-10. Mistral documents a voice-agent pipeline instead: Voxtral Realtime (streaming STT, Feb 2026) + an LLM + Voxtral TTS (Mar 2026), with open weights available.","status":"GA","models":[{"name":"Voxtral Realtime (Voxtral-Mini-4B-Realtime-2602)","status":"GA","notes":"Streaming STT, open weights (Apache-2.0); part of a cascade, not S2S."},{"name":"Voxtral TTS","status":"GA","notes":"Open-weight TTS (Mar 2026)."}],"transports":[],"audio":{"input":"n/a for S2S","output":"n/a for S2S"},"languages":"n/a","voices":"n/a","latency":"n/a","features":[],"pricing":{"model":"per-token","items":[],"audio_token_rate":"n/a","est_per_minute_usd":{"low":null,"high":null,"basis":"Cascade pricing only; see the STT/TTS segments."},"free_tier":"n/a","source":"https://docs.mistral.ai/studio-api/audio/overview"},"limits":[],"regions":"n/a","setup":{"steps":["Build a cascade (Voxtral Realtime -> LLM -> Voxtral TTS) or use another provider in this list."],"endpoint":"n/a","auth":"n/a","snippet_lang":"javascript","snippet":""},"warnings":[{"severity":"low","title":"Cascade, not native audio","detail":"Mistral's pipeline loses prosody and adds hops compared with native S2S models, but is self-hostable with open weights."}],"best_for":"Self-hosted or EU-sovereign cascaded voice agents (Mistral).","open_source":true,"self_hostable":true,"compliance":"n/a","docs":[{"label":"Mistral audio overview","url":"https://docs.mistral.ai/studio-api/audio/overview"}],"sources":["https://docs.mistral.ai/studio-api/audio/overview","https://www.gigazine.net/gsc_news/en/20260205-mistral-ai-voxtral-transcribe-2"],"confidence":"medium","unverified":"Meta's status is based on the absence of any official Llama API audio documentation in search results; Mistral prices and exact model ids were not verified for this segment.","cat":"voice"}],"providers":[{"id":"openai-realtime","name":"OpenAI Realtime API","vendor":"OpenAI","category":"speech-to-speech","summary":"Native speech-to-speech API for the gpt-realtime family over WebRTC, WebSocket or SIP, plus realtime transcription and live translation sessions. The default choice for production voice agents that need strong tool calling and reasoning.","status":"GA","models":[{"name":"gpt-realtime-2.1","status":"GA","notes":"Current flagship (released July 2026). Reasoning model with configurable reasoning effort, better alphanumerics, noise and interruption handling than gpt-realtime-2. 128k context, 32k max output. Text, audio and image input."},{"name":"gpt-realtime-2.1-mini","status":"GA","notes":"Distilled, cheaper reasoning model. Same 128k context. Replacement for gpt-realtime-mini."},{"name":"gpt-realtime-2","status":"GA","notes":"May 2026 release, same price as 2.1. Older but not yet on the deprecation list."},{"name":"gpt-realtime-1.5","status":"GA","notes":"Older generation; text output $16/1M instead of $24."},{"name":"gpt-realtime","status":"Deprecated","notes":"Shutdown January 20, 2027; replacement gpt-realtime-2.1."},{"name":"gpt-realtime-mini","status":"Deprecated","notes":"Shutdown January 20, 2027; replacement gpt-realtime-2.1-mini. The 2025-10-06 snapshot was already shut down July 23, 2026."},{"name":"gpt-realtime-translate","status":"GA","notes":"Speech-to-speech translation on /v1/realtime/translations, $0.034/min. Third-party reports say 70+ input and 13 output languages (not confirmed on the model page)."},{"name":"gpt-live-transcribe / gpt-realtime-whisper / gpt-transcribe","status":"GA","notes":"Realtime transcription session models (session.type = transcription). gpt-live-transcribe is the recommended one."},{"name":"gpt-4o-realtime-preview (all snapshots), gpt-4o-mini-realtime-preview","status":"Deprecated","notes":"Shut down May 7, 2026."}],"transports":["WebRTC","WebSocket","SIP"],"audio":{"input":"audio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk.","output":"audio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response."},"languages":"Multilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents.","voices":"10 built-in voices: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar. OpenAI recommends marin or cedar. No custom voices.","latency":"Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published.","features":["function calling (parallel tool calls)","remote MCP servers as tools","semantic_vad and server_vad turn detection, or manual push-to-talk (turn_detection null)","barge-in with conversation.item.truncate to sync what the user heard","image input (input_image) on gpt-realtime-2.x and gpt-realtime","configurable reasoning effort on 2.x models","input audio transcription in parallel","out-of-band responses (response.conversation = none)","automatic prompt caching","truncation controls (retention_ratio, token_limits.post_instructions)","SIP telephony with webhooks (realtime.call.incoming)","sideband server WebSocket for a WebRTC or SIP call (?call_id=)","realtime transcription sessions and live translation sessions","Agents SDK (@openai/agents/realtime) for browser and server"],"pricing":{"model":"per-token","items":[{"what":"gpt-realtime-2.1 audio input","price":"$32.00","unit":"per 1M tokens","notes":"cached $0.40"},{"what":"gpt-realtime-2.1 audio output","price":"$64.00","unit":"per 1M tokens","notes":""},{"what":"gpt-realtime-2.1 text input","price":"$4.00","unit":"per 1M tokens","notes":"cached $0.40"},{"what":"gpt-realtime-2.1 text output","price":"$24.00","unit":"per 1M tokens","notes":"includes reasoning tokens; higher reasoning effort raises output usage"},{"what":"gpt-realtime-2.1 image input","price":"$5.00","unit":"per 1M tokens","notes":"cached $0.50"},{"what":"gpt-realtime-2.1-mini audio input","price":"$10.00","unit":"per 1M tokens","notes":"cached $0.30"},{"what":"gpt-realtime-2.1-mini audio output","price":"$20.00","unit":"per 1M tokens","notes":""},{"what":"gpt-realtime-2.1-mini text input / output","price":"$0.60 / $2.40","unit":"per 1M tokens","notes":"cached input $0.06"},{"what":"gpt-realtime-2.1-mini image input","price":"$0.80","unit":"per 1M tokens","notes":"cached $0.08"},{"what":"gpt-realtime-2 (all lines)","price":"same as gpt-realtime-2.1","unit":"per 1M tokens","notes":""},{"what":"gpt-realtime-1.5 / gpt-realtime audio in / out","price":"$32.00 / $64.00","unit":"per 1M tokens","notes":"text $4.00 in, $16.00 out; cached $0.40"},{"what":"gpt-realtime-mini audio in / out","price":"$10.00 / $20.00","unit":"per 1M tokens","notes":"text $0.60 / $2.40"},{"what":"gpt-realtime-translate","price":"$0.034","unit":"per minute of audio","notes":"billed by duration"},{"what":"gpt-live-transcribe / gpt-realtime-whisper","price":"$0.017","unit":"per minute","notes":"realtime transcription sessions"},{"what":"gpt-transcribe","price":"$0.0045","unit":"per minute","notes":""},{"what":"gpt-4o-transcribe / gpt-4o-mini-transcribe (input transcription)","price":"$0.006 / $0.003","unit":"per minute (estimated)","notes":"token rates $2.50/$10.00 and $1.25/$5.00 per 1M"}],"audio_token_rate":"User audio 1 token per 100 ms = 10 tokens/s = 600 tokens/min. Assistant audio 1 token per 50 ms = 20 tokens/s = 1,200 tokens/min. VAD-filtered silence is not billed. (OpenAI realtime-costs guide)","est_per_minute_usd":{"low":0.096,"high":0.76,"basis":"Low = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached)."},"free_tier":"None. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow.","source":"https://developers.openai.com/api/docs/pricing"},"limits":["Max session length 60 minutes, no warning event before cutoff (third-party SDKs reconnect at ~50 min)","Context window 128,000 tokens; max output 32,000 (2.1 and 2.1-mini)","Default rate limits gpt-realtime-2.1: Build 400 RPM / 200,000 TPM, Launch 10,000 RPM / 4,000,000 TPM, Grow 20,000 RPM / 15,000,000 TPM","gpt-realtime-translate: Build 200, Launch 650, Grow 850 minutes of audio per minute","Ephemeral client secret lifetime 10 s to 7,200 s, default 600 s","Audio append chunk max 15 MB"],"regions":"Global API. Data residency in the United States and Europe (EEA + Switzerland) for gpt-realtime, -1.5, -mini, -2, -2.1, -2.1-mini; EU requires approved abuse-monitoring controls (ZDR, Modified Abuse Monitoring, etc). EU SIP endpoint sip-eu.api.openai.com. Tracing is not EU-residency compliant for /v1/realtime.","setup":{"steps":["Create an OpenAI Platform account and a project at platform.openai.com; add a payment method (free tier cannot use realtime).","Create a project API key and store it only on your server.","Browser: your server calls POST https://api.openai.com/v1/realtime/client_secrets with the session config and returns the ek_ value; the browser POSTs its SDP offer to https://api.openai.com/v1/realtime/calls with that key and opens the oai-events data channel.","Server or telephony bridge: open the WebSocket below with the API key and send session.update with session.type = realtime.","Phone: point your SIP trunk at sip:<project_id>@sip.api.openai.com;transport=tls, add a realtime.call.incoming webhook, then POST /v1/realtime/calls/{call_id}/accept.","Log response.done usage on every turn so you can see context growth and cache hit rates."],"endpoint":"wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1 (WebSocket); https://api.openai.com/v1/realtime/calls (WebRTC SDP); sip:$PROJECT_ID@sip.api.openai.com;transport=tls (SIP)","auth":"Authorization: Bearer <API key> for server WebSocket and SIP control. For browsers and mobile, mint an ephemeral key (ek_...) server side via POST /v1/realtime/client_secrets (expires_after.seconds 10 to 7200, default 600) and lock instructions/tools in that session config. Optional OpenAI-Safety-Identifier header on the server request is bound to the token.","snippet_lang":"javascript","snippet":"import WebSocket from \"ws\"; // server side only: never ship the API key to a browser\nconst ws = new WebSocket(\"wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1\", {\n  headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },\n});\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"session.update\",\n    session: {\n      type: \"realtime\",\n      instructions: \"You are a friendly support agent. Keep answers short.\",\n      audio: {\n        input: { format: { type: \"audio/pcm\", rate: 24000 }, turn_detection: { type: \"semantic_vad\" } },\n        output: { format: { type: \"audio/pcm\", rate: 24000 }, voice: \"marin\" },\n      },\n    },\n  }));\n});\n// Call for every ~100 ms chunk of 24 kHz mono PCM16 from your mic or phone bridge\nexport function sendAudio(pcm16) {\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: pcm16.toString(\"base64\") }));\n}\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"response.output_audio.delta\") playPcm16(Buffer.from(ev.delta, \"base64\"));\n  if (ev.type === \"input_audio_buffer.speech_started\") stopPlayback(); // barge-in: flush local audio\n  if (ev.type === \"response.done\") console.log(ev.response.usage); // log tokens per turn\n  if (ev.type === \"error\") console.error(ev.error);\n});"},"warnings":[{"severity":"high","title":"Context growth multiplies cost","detail":"Each response re-sends the whole conversation, including earlier audio (model audio is re-billed as audio input at $32/1M). A 10 minute call can cost 8x the naive per-minute figure with no cache hits. Changing instructions or tools mid-session, or truncating every turn, breaks the cache. Use token_limits.post_instructions with retention_ratio around 0.8, or delete or summarise old items."},{"severity":"high","title":"Beta interface is gone","detail":"The OpenAI-Beta: realtime=v1 interface was removed May 12, 2026. GA requires session.type, nests audio config under session.audio.input/output, and renames events (response.output_audio.delta, response.output_text.delta, response.output_audio_transcript.delta). Old tutorials and Azure preview samples using response.audio.delta will silently get no audio."},{"severity":"high","title":"Deprecation calendar","detail":"gpt-realtime, gpt-realtime-mini, gpt-4o-realtime and gpt-4o-mini-realtime shut down January 20, 2027. gpt-4o realtime previews already died May 7, 2026. Pin explicit model ids and test 2.1 before January."},{"severity":"medium","title":"Voice is locked after first audio","detail":"Once the model has emitted audio in a session the voice cannot be changed. Set it in the client secret or the first session.update."},{"severity":"medium","title":"60 minute hard cap without warning","detail":"Sessions end at 60 minutes and no warning event is sent. Track session age yourself and reconnect at a turn boundary, replaying a summary into the new session."},{"severity":"medium","title":"Reasoning effort trades latency and cost","detail":"gpt-realtime-2.x are reasoning models; higher reasoning effort increases both time to first audio and billed text output tokens ($24/1M). Keep effort low for chit-chat and raise it only for tool-heavy turns."},{"severity":"medium","title":"Mini models are less reliable with tools","detail":"OpenAI's own cost guide warns the mini model may follow instructions and call functions less reliably. Prototype on the full model, then measure regressions before switching."},{"severity":"low","title":"Translation endpoint has its own retention rules","detail":"/v1/realtime/translations is not listed in the ZDR column of OpenAI's data controls table, unlike /v1/realtime. Check with OpenAI before sending regulated audio."}],"best_for":"Production voice agents and phone bots that need the strongest tool calling and reasoning with a mature SDK, WebRTC in the browser and native SIP.","open_source":false,"self_hostable":false,"compliance":"/v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement.","docs":[{"label":"Realtime guide","url":"https://developers.openai.com/api/docs/guides/realtime"},{"label":"Realtime conversations (events, voices, VAD)","url":"https://developers.openai.com/api/docs/guides/realtime-conversations"},{"label":"Managing costs","url":"https://developers.openai.com/api/docs/guides/realtime-costs"},{"label":"WebRTC guide","url":"https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime"},{"label":"SIP guide","url":"https://developers.openai.com/api/docs/guides/voice-sip?api=realtime"},{"label":"Realtime transcription","url":"https://developers.openai.com/api/docs/guides/realtime-transcription"},{"label":"Deprecations","url":"https://developers.openai.com/api/docs/deprecations"}],"sources":["https://developers.openai.com/api/docs/pricing","https://developers.openai.com/api/docs/models/gpt-realtime-2.1","https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini","https://developers.openai.com/api/docs/models/gpt-realtime-translate","https://developers.openai.com/api/docs/guides/realtime-costs","https://developers.openai.com/api/docs/guides/realtime-conversations.md","https://developers.openai.com/api/docs/guides/voice-websockets?api=realtime","https://developers.openai.com/api/docs/guides/voice-webrtc?api=realtime","https://developers.openai.com/api/docs/guides/voice-sip?api=realtime","https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets/methods/create","https://developers.openai.com/api/docs/deprecations","https://developers.openai.com/api/docs/guides/your-data","https://datanorth.ai/news/openai-releases-gpt-realtime-2-1-voice-models"],"confidence":"high","unverified":"Exact reasoning-effort levels and how to set them for gpt-realtime-2.x; latency numbers (only the third-party-reported 25 percent p95 improvement); the language lists for gpt-realtime-translate; codecs accepted on inbound Realtime SIP; whether Realtime sessions have a per-tier concurrent session cap (model page lists RPM/TPM only).","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"OpenAI Realtime","facts":{"latency_ms":null,"languages":null,"max_session_min":60,"concurrency":null,"free_tier":false,"free_credit_usd":null,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":true,"websocket":true,"sip":true,"open_weights":false,"function_calling":true,"image_input":true,"voices":10,"voice_cloning":false,"context_tokens":128000,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":32,"price_audio_out_per_1m":64,"price_per_min":null,"_notes":"Prices are gpt-realtime-2.1 (mini: $10/$20). SOC 2 and BAA are account-level OpenAI programs; confirm realtime coverage. EU residency needs approved abuse-monitoring controls. Languages: multilingual, no published list. Latency: only a relative p95 claim."}},{"id":"openai-gpt-live","name":"OpenAI GPT-Live API","vendor":"OpenAI","category":"speech-to-speech","summary":"Full-duplex voice model (gpt-live-1) that listens while it speaks and hands reasoning and tools to a backend agent, billed per minute. Same voice system as ChatGPT Voice; suits natural, interruption-heavy conversations where you already have an agent backend.","status":"GA","models":[{"name":"gpt-live-1","status":"GA","notes":"API launch reported Sept 10, 2026 (third-party coverage); model page does not label preview or GA. Audio and text in and out, no image or video. Knowledge cutoff Jul 31, 2025. Delegates to a Responses model (e.g. gpt-5.6-terra or gpt-5.6-luna) or to your own backend."}],"transports":["WebRTC","WebSocket","SIP"],"audio":{"input":"WebSocket: audio/pcm 24 kHz (default) or 16 kHz mono 16-bit LE, audio/pcmu or audio/pcma 8 kHz; raw bytes base64, even byte length, no container. Format fixed at session start. WebRTC and SIP negotiate codecs (outbound SIP needs Opus + SDES-SRTP).","output":"Same format as input (one setting covers both)."},"languages":"Not listed on the model page.","voices":"12 new voices listed in the docs (quartz, ripple, vesper, willow, stone, gleam, meridian, bossa, tempo, beacon, delta, cinder); docs give marin as the default. Voice change requires a new session. Audio is SynthID watermarked (reported July 31, 2026 update).","latency":"Vendor claim: improves Full Duplex Bench score by 30 percentage points over gpt-realtime-2.1 (reported via third-party coverage). No millisecond figure published.","features":["full duplex (listens while speaking), smooth interruption handling","delegation to a Responses backend managed by OpenAI, or client delegation to your own agent","function calling and backend tools (e.g. web_search)","input and output transcript deltas","context appends during the session (up to 500 tokens per event)","startup history seeding (up to 128 messages / 8,192 tokens)","session forking from a stored recording (not under ZDR)","sideband server connection to monitor or control a session","inbound and outbound SIP calling (outbound must be enabled for your org)","partner integrations: LiveKit, Twilio, Telnyx, Daily/Pipecat"],"pricing":{"model":"per-minute","items":[{"what":"gpt-live-1 voice session","price":"$0.05","unit":"per minute, billed per second (not rounded up)","notes":"covers the voice layer only"},{"what":"Backend model and tools (Responses delegation)","price":"standard model rates","unit":"per token / per tool call","notes":"billed separately at the normal price of the chosen backend model and tools"}],"audio_token_rate":"Not applicable for the voice layer (duration billing). Backend usage is ordinary text tokens.","est_per_minute_usd":{"low":0.05,"high":null,"basis":"Low = voice layer only, $0.05 per minute of session duration with no delegated work. High depends entirely on the backend model, how often the model delegates and tool fees; budget per-minute session cost plus backend token spend measured in your own tests."},"free_tier":"Free tier not supported.","source":"https://developers.openai.com/api/docs/models/gpt-live-1"},"limits":["Concurrent sessions (OpenAI direct): Build 50, Launch 300, Grow 500; Free unsupported","Context window 128,000 tokens; above 90 percent usage a replacement engine starts with up to 8,192 tokens of history, so older details may be summarised or dropped","Instructions up to 16,384 tokens; startup history up to 128 messages / 8,192 tokens","Session ends with reason expired at a duration limit that the docs do not state","Outbound SIP: 3 minutes ringing, 2 hours max connected call, 1 MiB request body (not configurable)","Stored recordings kept 30 days when store: true"],"regions":"Global API; data residency in the United States and Europe for gpt-live-1. Also offered on Azure (see Azure entry) with its own concurrent-session tiers.","setup":{"steps":["Use a paid OpenAI project (free tier unsupported) and create a project API key kept on your server.","Browser: create an RTCPeerConnection with the mic track and an oai-events data channel, send the SDP offer to your own server.","Server: POST /v1/live/sessions with session.model = gpt-live-1, delegation settings and transport { type: webrtc, sdp }; return the SDP answer from the 201 response.","Wait for session.started on the data channel before speaking; do not send session.start over WebRTC.","Server-side audio: open wss://api.openai.com/v1/live/sessions and send session.start as the first message.","End with session.close and wait for session.closed to collect final usage."],"endpoint":"POST https://api.openai.com/v1/live/sessions (WebRTC session create); wss://api.openai.com/v1/live/sessions (WebSocket); SIP via live.transport.incoming webhook and POST /v1/live/sessions/{session_id}/accept","auth":"Project API key on your server. The docs do not document an ephemeral client-secret flow for Live: keep the browser talking to your server for the SDP exchange.","snippet_lang":"javascript","snippet":"import WebSocket from \"ws\"; // server side; browsers should use WebRTC via your server\nconst ws = new WebSocket(\"wss://api.openai.com/v1/live/sessions\", {\n  headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },\n});\nws.on(\"open\", () => ws.send(JSON.stringify({\n  type: \"session.start\",\n  session: {\n    model: \"gpt-live-1\",\n    instructions: \"Be concise. Delegate anything that needs current information.\",\n    audio: { format: { type: \"audio/pcm\", rate: 24000 }, output: { voice: \"marin\" } },\n    delegation: {\n      type: \"responses\",\n      responses: { model: \"gpt-5.6-luna\", tools: [{ type: \"web_search\" }], tool_choice: \"auto\" },\n    },\n  },\n})));\nlet ready = false;\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"session.started\") ready = true; // do not send audio before this\n  if (ev.type === \"session.output_audio.delta\") playPcm16(Buffer.from(ev.delta, \"base64\"));\n  if (ev.type === \"session.closed\") console.log(\"final usage\", ev);\n  if (ev.type === \"error\") console.error(ev.error);\n});\n// 24 kHz mono PCM16, even byte length, no WAV header\nexport function sendAudio(pcm16) {\n  if (ready) ws.send(JSON.stringify({ type: \"session.input_audio.append\", audio: pcm16.toString(\"base64\") }));\n}\n// To end: ws.send(JSON.stringify({ type: \"session.close\" })) and wait for \"session.closed\""},"warnings":[{"severity":"high","title":"Two bills, not one","detail":"$0.05/min covers only the voice layer. Every delegated task runs a Responses model plus tools billed at normal rates, which can easily exceed the voice cost on tool-heavy calls. Log backend usage per session."},{"severity":"high","title":"Billed by wall-clock duration","detail":"Billing is per second of session duration, so idle or forgotten sessions keep costing. Always send session.close and set your own inactivity timeout."},{"severity":"medium","title":"Long calls lose memory","detail":"When context passes 90 percent of 128k, GPT-Live swaps to a replacement engine with at most 8,192 tokens of history. Keep critical facts in instructions or re-append them."},{"severity":"medium","title":"Different API from Realtime","detail":"GPT-Live uses /v1/live/sessions with session.start, session.input_audio.append and session.output_audio.delta. Realtime code and events do not carry over; follow the migration guide. Delegation mode cannot be changed mid-session."},{"severity":"medium","title":"Transcript events need stitching","detail":"Transcript deltas have no item id or turn-complete event and output audio deltas have no timing fields, so your app must group them itself."},{"severity":"medium","title":"Moderation can cut audio","detail":"Moderation can stop assistant audio without ending the session, or end the session. Handle error events with client_event_id and show the user something."},{"severity":"low","title":"Outbound calls are not idempotent","detail":"Each outbound create places a new call and X-Client-Request-Id does not deduplicate. Never auto-retry after an ambiguous timeout."},{"severity":"low","title":"Recordings and ZDR","detail":"store: true keeps recordings 30 days; under ZDR store is forced false and forking or recording download is unavailable."}],"best_for":"Natural, full-duplex conversations (companions, concierge, phone agents) where an existing agent or Responses model does the thinking and you want simple per-minute voice pricing.","open_source":false,"self_hostable":false,"compliance":"/v1/live/sessions is ZDR eligible with limitations (store forced false, no forking or recording download). Abuse-monitoring logs 30 days. US and EU data residency.","docs":[{"label":"GPT-Live guide","url":"https://developers.openai.com/api/docs/guides/live"},{"label":"Managing sessions","url":"https://developers.openai.com/api/docs/guides/live-conversations"},{"label":"WebRTC quickstart","url":"https://developers.openai.com/api/docs/guides/voice-webrtc?api=live"},{"label":"WebSocket guide","url":"https://developers.openai.com/api/docs/guides/voice-websockets?api=live"},{"label":"Model page","url":"https://developers.openai.com/api/docs/models/gpt-live-1"}],"sources":["https://developers.openai.com/api/docs/models/gpt-live-1","https://developers.openai.com/api/docs/guides/live","https://developers.openai.com/api/docs/guides/live-conversations","https://developers.openai.com/api/docs/guides/voice-webrtc?api=live","https://developers.openai.com/api/docs/guides/voice-websockets?api=live","https://developers.openai.com/api/docs/guides/voice-sip?api=realtime","https://developers.openai.com/api/docs/guides/your-data","https://windowsreport.com/gpt-live-1-is-now-available-to-developers-through-the-openai-api/","https://unite.ai/openais-gpt-live-1-arrives-in-the-api-at-0-05-per-minute"],"confidence":"medium","unverified":"Formal GA/preview label (model page silent); maximum session duration; supported languages; voice list vs. default (docs list 12 names but give marin as default); launch date (Sept 10, 2026 from secondary sources); latency figures. The Free-tier/concurrency table differs slightly between search snippets (Tier 1-5) and the current model page (Build/Launch/Grow).","cat":"voice","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":50,"free_tier":false,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":true,"websocket":true,"sip":true,"open_weights":false,"function_calling":true,"image_input":false,"voices":12,"voice_cloning":null,"context_tokens":128000,"byo_llm":true,"native_s2s":true,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.05,"_notes":"Concurrency is the Build tier (Launch 300, Grow 500). $0.05/min covers the voice layer only; delegated backend model and tools billed separately. Can delegate to your own agent backend. Session has an unstated duration limit; outbound SIP calls cap at 2 hours."}},{"id":"azure-openai-realtime","name":"Azure OpenAI GPT Realtime API (Microsoft Foundry)","vendor":"Microsoft","category":"speech-to-speech","summary":"The OpenAI gpt-realtime models hosted in your Azure subscription, with Entra ID auth, Azure networking and Data Zone options. Suits teams that must stay inside Azure contracts or regions.","status":"GA","models":[{"name":"gpt-realtime-2.1 (2026-07-07)","status":"GA","notes":"Global and Data Zone deployments."},{"name":"gpt-realtime-2.1-mini (2026-07-07)","status":"GA","notes":"Global and Data Zone."},{"name":"gpt-realtime-2 (2026-05-07)","status":"GA","notes":"Global."},{"name":"gpt-realtime-1.5 (2026-02-23)","status":"GA","notes":"Global and Data Zone."},{"name":"gpt-realtime (2025-08-28), gpt-realtime-mini (2025-10-06, 2025-12-15)","status":"GA","notes":"Older; OpenAI retires these Jan 20, 2027, check the Azure retirement schedule."},{"name":"gpt-4o-realtime-preview / gpt-4o-mini-realtime-preview (2024-12-17)","status":"Preview","notes":"Still listed on Azure; preview endpoints are deprecated."},{"name":"gpt-realtime-translate, gpt-realtime-whisper (2026-05-06), gpt-live-transcribe (2026-07-29)","status":"GA","notes":"Duration-billed translation and transcription."},{"name":"gpt-live-1 (GPT-Live)","status":"Preview","notes":"Listed in Azure quota docs with its own concurrent session tiers; availability status on Azure not confirmed."}],"transports":["WebRTC","WebSocket","SIP"],"audio":{"input":"PCM16 mono 24 kHz recommended (send ~100 ms chunks); G.711 supported per the shared OpenAI event model; WebRTC negotiates codecs.","output":"PCM16 24 kHz (same options as OpenAI)."},"languages":"Same models as OpenAI; Microsoft advises validating languages with production-like audio and passing ISO-639-1 hints for transcription.","voices":"Same 10 OpenAI voices (marin and cedar recommended). For Azure neural or custom voices use Voice Live instead.","latency":"Microsoft guidance (transport only, not model time): WebRTC ~100 ms, WebSocket ~200 ms.","features":["function calling and remote MCP servers","server_vad, semantic_vad or manual turn handling","image input","out-of-band responses","Entra ID keyless auth and managed identity","ephemeral tokens via /openai/v1/realtime/client_secrets","SIP telephony","Global, Data Zone (US/EU) deployment types"],"pricing":{"model":"per-token","items":[{"what":"gpt-realtime-2.1 Global audio input / output","price":"$32.00 / $64.00","unit":"per 1M tokens","notes":"cached audio input $0.40"},{"what":"gpt-realtime-2.1 Global text input / output","price":"$4.00 / $24.00","unit":"per 1M tokens","notes":"cached $0.40"},{"what":"gpt-realtime-2.1 Global image input","price":"$5.00","unit":"per 1M tokens","notes":"cached $0.50"},{"what":"gpt-realtime-2.1 Data Zone audio input / output","price":"$35.20 / $70.40","unit":"per 1M tokens","notes":"cached $0.44; text $4.40 / $26.40; image $5.50"},{"what":"gpt-realtime-2.1-mini Global audio input / output","price":"$10.00 / $20.00","unit":"per 1M tokens","notes":"cached $0.30; text $0.60 / $2.40; image $0.80"},{"what":"gpt-realtime-2.1-mini Data Zone audio input / output","price":"$11.00 / $22.00","unit":"per 1M tokens","notes":"cached $0.33; text $0.66 / $2.64"},{"what":"gpt-realtime-2 Global","price":"$32.00 / $64.00 audio, $4.00 / $24.00 text","unit":"per 1M tokens","notes":"cached $0.40"},{"what":"gpt-realtime-1.5 Global / Data Zone audio","price":"$32.00 / $64.00 and $35.20 / $70.40","unit":"per 1M tokens","notes":"text out $16 Global, $17.60 DZ"},{"what":"gpt-realtime-translate (Global)","price":"$2.04","unit":"per hour","notes":"= $0.034/min"},{"what":"gpt-realtime-whisper (Global)","price":"$1.02","unit":"per hour","notes":"= $0.017/min"}],"audio_token_rate":"Same as OpenAI: about 10 tokens/s input audio, 20 tokens/s output audio (Microsoft Voice Live token table for Azure OpenAI models).","est_per_minute_usd":{"low":0.096,"high":0.84,"basis":"Low = gpt-realtime-2.1 Global, 1 min user audio (600 tokens) + 1 min model audio (1,200 tokens), single turn, no cache. Data Zone single turn $0.106. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. High shown at Data Zone rates ($0.76 Global); with full cache hits about $0.06/min."},"free_tier":"No free tier for realtime models (Azure free account credits can apply).","source":"https://prices.azure.com/api/retail/prices (Azure Retail Prices API, eastus2 meters, retrieved 2026-10-10); https://azure.microsoft.com/pricing/details/cognitive-services/openai-service/"},"limits":["Max session duration 60 minutes (monitor expires_at in session.created)","Default gpt-realtime quota in the Tier 1 table: 200 RPM and 100,000 TPM (GlobalStandard); higher tiers raise it (e.g. 300 RPM / 150,000 TPM)","GPT-Live on Azure: concurrent sessions per subscription Default 10, Tier 1 25, Tier 2 50, Tier 3 200, Tier 4 300, Tier 5 500","Realtime quota is separate from chat completions quota","Doc still states 32,000 input / 4,096 output tokens for the Realtime API, which conflicts with the 128k model context on OpenAI"],"regions":"Global deployments served from East US 2 and Sweden Central for WebRTC/realtime; Data Zone (US, EU) keeps processing inside the zone. Check the region availability page for each model.","setup":{"steps":["Create an Azure subscription and a Microsoft Foundry resource in a supported region (East US 2 or Sweden Central are the documented realtime regions).","In the Foundry portal deploy a realtime model (e.g. gpt-realtime-2.1) as Global or Data Zone; note the deployment name.","Prefer Entra ID: assign Cognitive Services OpenAI User and request tokens for scope https://ai.azure.com/.default; otherwise copy the resource key.","Server or telephony: open the GA WebSocket (/openai/v1/realtime) with the deployment name as model. Do not add api-version.","Browser: your token service calls https://<resource>.openai.azure.com/openai/v1/realtime/client_secrets and the browser posts its SDP to https://<resource>.openai.azure.com/openai/v1/realtime/calls with the ephemeral token.","Request quota increases early: 100k TPM is small for realtime because of context re-billing."],"endpoint":"wss://<resource>.openai.azure.com/openai/v1/realtime?model=<deployment> (WebSocket); https://<resource>.openai.azure.com/openai/v1/realtime/calls (WebRTC); https://<resource>.openai.azure.com/openai/v1/realtime/client_secrets (ephemeral tokens)","auth":"api-key header or Authorization: Bearer <Entra ID token> on the server. Browsers get an ephemeral token from your token service; the more secure pattern proxies the SDP exchange so the browser never holds even the ephemeral token.","snippet_lang":"javascript","snippet":"import WebSocket from \"ws\"; // server side; browsers use /openai/v1/realtime/client_secrets + WebRTC\nconst RESOURCE = process.env.AZURE_OPENAI_RESOURCE;      // e.g. my-foundry-resource\nconst DEPLOYMENT = process.env.AZURE_OPENAI_DEPLOYMENT;  // your deployment name, not the model name\nconst ws = new WebSocket(\n  `wss://${RESOURCE}.openai.azure.com/openai/v1/realtime?model=${DEPLOYMENT}`,\n  { headers: { \"api-key\": process.env.AZURE_OPENAI_API_KEY } } // or Authorization: Bearer <Entra token>\n);\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"session.update\",\n    session: {\n      type: \"realtime\",\n      instructions: \"You are a friendly support agent. Keep answers short.\",\n      audio: {\n        input: { format: { type: \"audio/pcm\", rate: 24000 }, turn_detection: { type: \"server_vad\" } },\n        output: { format: { type: \"audio/pcm\", rate: 24000 }, voice: \"marin\" },\n      },\n    },\n  }));\n});\nexport function sendAudio(pcm16) { // 24 kHz mono PCM16, ~100 ms chunks\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: pcm16.toString(\"base64\") }));\n}\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"response.output_audio.delta\") playPcm16(Buffer.from(ev.delta, \"base64\"));\n  if (ev.type === \"input_audio_buffer.speech_started\") stopPlayback();\n  if (ev.type === \"error\") console.error(ev.error);\n});"},"warnings":[{"severity":"high","title":"Preview endpoints and samples are deprecated","detail":"Old samples use /openai/realtimeapi/sessions, https://<region>.realtimeapi-preview.ai.azure.com/v1/realtimertc, api-version query strings and beta event names (response.audio.delta). The GA path is /openai/v1/realtime/... with no api-version; mixing them gives 404s or silent audio."},{"severity":"high","title":"Low default TPM","detail":"The Tier 1 table shows gpt-realtime at 100,000 TPM. Because every turn re-sends history, a handful of concurrent long calls can hit 429s. Plan quota requests and back-off before launch."},{"severity":"medium","title":"Data Zone costs 10 percent more","detail":"Data Zone meters are exactly 1.1x Global (e.g. $35.20 vs $32.00 audio input). Use Data Zone only when residency requires it."},{"severity":"medium","title":"Model is a deployment, not a name","detail":"The model query parameter is your deployment name. Upgrading to a new model version means a new or updated deployment, and retired versions follow Azure's own retirement schedule, not OpenAI's."},{"severity":"medium","title":"Entra auth gotcha","detail":"Keyless auth fails if the AZURE_OPENAI_API_KEY environment variable is set; the SDK picks the key. Unset it when using DefaultAzureCredential."},{"severity":"medium","title":"Region-limited availability","detail":"Realtime Global deployments are documented for East US 2 and Sweden Central only. Latency from other continents can be noticeably worse than OpenAI's global edge."},{"severity":"low","title":"Inconsistent limits in docs","detail":"The how-to page states both a 60 minute session cap and a 32,000 input / 4,096 output token limit that predates 128k-context models. Test the real limits on your deployment."}],"best_for":"Enterprises already on Azure that need private networking, Entra ID, Microsoft contracts and US/EU Data Zone processing for OpenAI realtime voice.","open_source":false,"self_hostable":false,"compliance":"Covered by Azure OpenAI enterprise terms (Microsoft Products and Services DPA; HIPAA BAA via Microsoft for in-scope Azure services). Data Zone deployments keep processing within the US or EU zone. Content filtering applies.","docs":[{"label":"GPT Realtime API how-to","url":"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio"},{"label":"Realtime via WebRTC","url":"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio-webrtc"},{"label":"Realtime via WebSockets","url":"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio-websockets"},{"label":"Quotas and limits","url":"https://learn.microsoft.com/en-us/azure/foundry/openai/quotas-limits"},{"label":"Pricing","url":"https://azure.microsoft.com/pricing/details/cognitive-services/openai-service/"}],"sources":["https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio","https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio-webrtc","https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio-websockets","https://learn.microsoft.com/en-us/azure/foundry/openai/quotas-limits","https://prices.azure.com/api/retail/prices","https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live"],"confidence":"medium","unverified":"Full raw WebSocket URL (docs show only the SDK base /openai/v1 with model=<deployment>; the ?model= form is inferred); Azure SIP URI format; per-model region list beyond East US 2 / Sweden Central; whether gpt-live-1 is generally available on Azure; retail prices are from the eastus2 Retail Prices API because the public pricing page renders prices client-side.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Azure OpenAI Realtime","facts":{"latency_ms":null,"languages":null,"max_session_min":60,"concurrency":null,"free_tier":false,"free_credit_usd":null,"hipaa":true,"soc2":null,"gdpr_eu":true,"webrtc":true,"websocket":true,"sip":true,"open_weights":false,"function_calling":true,"image_input":true,"voices":10,"voice_cloning":false,"context_tokens":128000,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":32,"price_audio_out_per_1m":64,"price_per_min":null,"_notes":"Global prices for gpt-realtime-2.1; Data Zone is 1.1x ($35.20/$70.40). Microsoft quotes transport latency only (WebRTC ~100 ms, WebSocket ~200 ms). Docs still state a 32k input token limit that conflicts with the 128k model context. Generic Azure free-account credits may apply."}},{"id":"azure-voice-live","name":"Azure Speech Voice Live API","vendor":"Microsoft","category":"speech-to-speech","summary":"Managed voice-agent API in Microsoft Foundry that wraps either native speech-to-speech models (gpt-realtime family, azure-realtime, phi4-mm-realtime) or text LLMs with Azure speech-to-text and 600+ Azure TTS voices, adding noise suppression, server echo cancellation, semantic VAD and avatars. Suits contact centers and branded agents on Azure.","status":"GA","models":[{"name":"gpt-realtime-2.1 / -datazone / -regional","status":"GA","notes":"Pro tier. Native audio, optional Azure TTS or custom voice output."},{"name":"gpt-realtime-2.1-mini, gpt-realtime-mini","status":"GA","notes":"Standard tier."},{"name":"gpt-realtime, gpt-realtime-1.5 (+ datazone/regional)","status":"GA","notes":"Pro tier."},{"name":"azure-realtime","status":"Preview","notes":"Microsoft's own dedicated realtime model with 34 azure-realtime-native voices (default ava). Requires API version 2026-01-01-preview or later. Pro tier."},{"name":"phi4-mm-realtime","status":"Preview","notes":"Lite tier. Phi-4 multimodal audio input + Azure TTS output."},{"name":"gpt-5.6-terra, gpt-5.4, gpt-5.2, gpt-5.1, gpt-5, gpt-4.1, gpt-4o","status":"GA","notes":"Pro tier, cascaded: Azure STT -> LLM -> Azure TTS."},{"name":"gpt-5.6-luna, gpt-5-mini, gpt-4.1-mini, gpt-4o-mini","status":"GA","notes":"Standard tier, cascaded."},{"name":"gpt-5-nano, gpt-4.1-nano","status":"GA","notes":"Lite tier, cascaded."},{"name":"BYOM (gpt-5.5, gpt-5.4-mini, gpt-5.4-nano and other Foundry deployments)","status":"Preview","notes":"Bring your own Foundry model deployment."}],"transports":["WebSocket","WebRTC"],"audio":{"input":"PCM16 mono at 24 kHz (default) or 16 kHz via input_audio_sampling_rate; Live-Reference AEC mode takes interleaved stereo PCM16 (mic + playback reference). Fixed for the session.","output":"PCM16 audio from the native model or Azure TTS; word timestamps and visemes available with Azure voices; avatar video over WebRTC (H.264)."},"languages":"146 input locales and 151 output locales per the FAQ (Azure speech); azure_semantic_vad_multilingual covers English, Spanish, French, Italian, German, Japanese, Portuguese, Chinese, Korean, Hindi.","voices":"600+ Azure neural voices across 150+ locales, 30+ HD (DragonHD) voices, MAI-Voice-2-Flash (preview), 34 azure-realtime-native voices, native gpt-realtime voices, and custom voice (limited access).","latency":"No numeric vendor claim found; marketed as low-latency. Text-LLM cascades are described by Microsoft as having slightly higher latency than native realtime models.","features":["azure_semantic_vad and azure_semantic_vad_multilingual turn detection (with filler-word removal), plus server_vad and semantic_vad","azure_deep_noise_suppression","server_echo_cancellation, plus Live-Reference AEC (client-supplied playback reference, API 2026-07-15+)","function calling incl. asynchronous calls; MCP in model mode (not with phi models)","Foundry Agent Service integration (agent_id + project_id)","phrase lists and custom speech models for recognition","custom voice, HD voice temperature, speaking rate, custom lexicon","word-level audio timestamps and visemes","text-to-speech avatars, photo avatars (VASA-1, VASA-2 preview), custom avatars","transcription models: azure-speech, mai-transcribe-2 (preview), whisper-1, gpt-4o-transcribe(-diarize)","telephony through Azure Communication Services"],"pricing":{"model":"per-token","items":[{"what":"Pro: LLM audio input / output (e.g. gpt-realtime-2.1 native audio)","price":"$32.00 / $64.00","unit":"per 1M tokens","notes":"cached audio $0.40"},{"what":"Pro: LLM text input / output","price":"$4.00 / $16.00 (a second Pro text-output meter is $24.00)","unit":"per 1M tokens","notes":"cached $0.40; image input $5.00, cached $0.50"},{"what":"Pro: Azure standard speech audio input / output","price":"$17.00 / $31.00","unit":"per 1M tokens","notes":"cached $0.40"},{"what":"Pro: Azure custom speech audio input / output","price":"$40.00 / $55.00","unit":"per 1M tokens","notes":"cached $0.40"},{"what":"Standard: LLM audio input / output (e.g. gpt-realtime-2.1-mini)","price":"$11.00 / $22.00","unit":"per 1M tokens","notes":"cached $0.33"},{"what":"Standard: LLM text input / output","price":"$0.66 / $2.64","unit":"per 1M tokens","notes":"cached $0.33"},{"what":"Standard: Azure standard speech audio input / output","price":"$15.00 / $26.00","unit":"per 1M tokens","notes":"cached $0.33"},{"what":"Standard: Azure custom speech audio input / output","price":"$39.00 / $50.00","unit":"per 1M tokens","notes":""},{"what":"Lite: LLM audio input / output (phi4-mm-realtime)","price":"$4.00 / $12.00","unit":"per 1M tokens","notes":"cached $0.04"},{"what":"Lite: LLM text input / output","price":"$0.11 / $0.44","unit":"per 1M tokens","notes":"cached $0.04"},{"what":"Lite: Azure standard speech audio input / output","price":"$15.00 / $25.00","unit":"per 1M tokens","notes":""},{"what":"Lite: Azure custom speech audio input / output","price":"$38.00 / $50.00","unit":"per 1M tokens","notes":""},{"what":"BYO model: standard speech audio input / output","price":"$12.50 / $23.00","unit":"per 1M tokens","notes":"custom speech $36.00 / $47.00"},{"what":"Avatars, custom voice training and hosting","price":"separate","unit":"Speech service pricing","notes":"e.g. TTS standard avatar realtime $0.50/min, HD standard avatar $0.70/min (Retail Prices API meters)"}],"audio_token_rate":"Microsoft: Azure OpenAI models ~10 tokens/s input audio and ~20 tokens/s output audio; Phi models ~12.5 tokens/s input, ~20 tokens/s output. Token rate for Azure speech (STT/TTS) meters is not published.","est_per_minute_usd":{"low":0.017,"high":0.76,"basis":"Low = Lite tier phi4-mm-realtime, 1 min user audio (750 tokens x $4/1M) + 1 min model audio (1,200 tokens x $12/1M), single turn. Pro tier with gpt-realtime-2.1 native audio matches Azure OpenAI Global: $0.096 single turn. High = Pro gpt-realtime-2.1, High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Standard tier gpt-realtime-2.1-mini: $0.033 low, $0.26 high. Avatar minutes and custom voice are extra."},"free_tier":"No dedicated free tier found for Voice Live (Azure Speech F0 is limited to one concurrent request).","source":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live#pricing and https://prices.azure.com/api/retail/prices (Voice Live meters, eastus2, retrieved 2026-10-10)"},"limits":["100,000 tokens per minute per resource by default (increase on request)","Max session duration and concurrency are not stated in Voice Live docs","Phrase list under 500 words/phrases","SIP is not supported directly; use Azure Communication Services for telephony","HD voices only in southeastasia, centralindia, swedencentral, westeurope, eastus, eastus2, westus2 (Voice Live routes synthesis to a supported region)","Content filtering cannot be modified or disabled (use BYOM for custom filtering)"],"regions":"10+ Azure regions; LLM availability and processing scope (global, data zone, regional) depend on the resource region and the model suffix.","setup":{"steps":["Create an Azure subscription and a Microsoft Foundry resource (recommended over a plain Speech resource; Speech resources lack Agent Service and BYOM).","Assign Cognitive Services User and Foundry User roles to your identity, or copy the resource key.","Pick a model by name (it sets the Pro / Standard / Lite price); no deployment is needed for natively supported models.","Connect a server WebSocket to the voice-live/realtime endpoint with api-version and model query parameters.","Send session.update with turn detection, noise suppression, echo cancellation and voice, then stream input_audio_buffer.append.","For browsers use the Voice Live WebRTC flow or relay through your backend; for avatars complete the session.avatar.connect SDP exchange."],"endpoint":"wss://<foundry-resource>.services.ai.azure.com/voice-live/realtime?api-version=2026-04-10&model=gpt-realtime-2.1 (older resources: <name>.cognitiveservices.azure.com)","auth":"Recommended Microsoft Entra token (scope https://ai.azure.com/.default) as Authorization: Bearer. API key works as an api-key header (server only) or an api-key query parameter; never use the query-string key in a browser.","snippet_lang":"javascript","snippet":"import WebSocket from \"ws\"; // server side (API key in header is not possible from a browser)\nconst RES = process.env.FOUNDRY_RESOURCE; // Foundry resource name\nconst url = `wss://${RES}.services.ai.azure.com/voice-live/realtime?api-version=2026-04-10&model=gpt-realtime-2.1`;\nconst ws = new WebSocket(url, { headers: { \"api-key\": process.env.FOUNDRY_API_KEY } }); // or Authorization: Bearer <Entra token>\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"session.update\",\n    session: {\n      instructions: \"You are a friendly support agent. Keep answers short.\",\n      input_audio_sampling_rate: 24000,\n      turn_detection: { type: \"azure_semantic_vad\", silence_duration_ms: 500, remove_filler_words: true },\n      input_audio_noise_reduction: { type: \"azure_deep_noise_suppression\" },\n      input_audio_echo_cancellation: { type: \"server_echo_cancellation\" },\n      voice: { name: \"en-US-Ava:DragonHDLatestNeural\", type: \"azure-standard\" },\n    },\n  }));\n});\nexport function sendAudio(pcm16) { // mono PCM16 at 24 kHz (or 16 kHz if configured)\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: pcm16.toString(\"base64\") }));\n}\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  // Voice Live follows the Azure OpenAI realtime event set; accept both naming styles\n  if (ev.type === \"response.audio.delta\" || ev.type === \"response.output_audio.delta\")\n    playPcm16(Buffer.from(ev.delta, \"base64\"));\n  if (ev.type === \"input_audio_buffer.speech_started\") stopPlayback();\n  if (ev.type === \"error\") console.error(ev.error);\n});"},"warnings":[{"severity":"high","title":"You pay for speech twice in cascaded mode","detail":"With text LLMs (gpt-5.x, gpt-4.1) you pay LLM text tokens plus Azure speech input and output audio tokens ($15 to $17 in, $25 to $31 out per 1M on standard voices, more for custom). Microsoft does not publish the tokens-per-second rate for Azure speech meters, so measure usage events in a pilot before quoting customers."},{"severity":"medium","title":"Price tier follows the model","detail":"You do not choose Pro, Standard or Lite; the model name decides. Mixing a Lite model with a custom voice bills the custom voice at the Pro rate."},{"severity":"medium","title":"Event names differ from OpenAI GA","detail":"Voice Live follows the Azure OpenAI realtime reference, whose examples still use flat fields (input_audio_format, voice string or object) and response.audio.delta. Do not paste OpenAI GA session.audio.* configs without checking the Voice Live reference."},{"severity":"medium","title":"Echo cancellation timing","detail":"Default server echo cancellation assumes you play audio as soon as it arrives; more than 2 s playback delay degrades it. Use Live-Reference AEC (API 2026-07-15+) if you buffer, mix or resample audio on the client."},{"severity":"medium","title":"Low default TPM","detail":"100k tokens per minute per resource is shared by every session on that resource. Long calls re-send history, so request an increase before load testing."},{"severity":"medium","title":"No direct SIP","detail":"Phone calls need Azure Communication Services (or another media bridge) in front of Voice Live."},{"severity":"low","title":"Limited-access features","detail":"Custom voice and custom avatar need an approved intake form; plan weeks of lead time."},{"severity":"low","title":"Preview pieces","detail":"azure-realtime, phi4-mm-realtime, MAI Transcribe, MAI-Voice-2-Flash, VASA-2 photo avatars and BYOM are preview and can change."}],"best_for":"Contact centers and branded assistants on Azure that want noise suppression, server echo cancellation, multilingual semantic VAD, hundreds of voices, custom voice or a talking avatar without stitching services together.","open_source":false,"self_hostable":false,"compliance":"Azure AI services enterprise terms; content filtering always on; data zone and regional model variants for residency. HIPAA BAA via Microsoft for in-scope Azure services (confirm Voice Live scope with Microsoft).","docs":[{"label":"Voice Live overview and pricing","url":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live"},{"label":"How to use Voice Live","url":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-how-to"},{"label":"FAQ","url":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-faq"},{"label":"Samples","url":"https://github.com/microsoft-foundry/voicelive-samples"}],"sources":["https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live","https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-how-to","https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-faq","https://prices.azure.com/api/retail/prices"],"confidence":"medium","unverified":"Session max duration and concurrency; tokens-per-second for Azure speech input/output meters; which Pro text-output meter ($16 vs $24) applies to which model (likely $24 for gpt-realtime-2.1); retail prices taken from eastus2 meters and may differ by region; WebRTC endpoint details for Voice Live.","cat":"voice","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":146,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"hipaa":true,"soc2":null,"gdpr_eu":true,"webrtc":true,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":true,"voices":600,"voice_cloning":true,"context_tokens":null,"byo_llm":true,"native_s2s":true,"price_audio_in_per_1m":32,"price_audio_out_per_1m":64,"price_per_min":null,"_notes":"Supports both native speech-to-speech models and a cascaded STT+LLM+TTS mode. 146 input / 151 output locales. 600+ Azure neural voices; custom voice is limited access. Prices are Pro tier with gpt-realtime-2.1 (Lite from $4/$12). HIPAA via Microsoft BAA for in-scope services; confirm Voice Live scope. Telephony via Azure Communication Services, not direct SIP. BYO model is preview."}},{"id":"google-gemini-live-api","name":"Gemini Live API (Gemini Developer API / Google AI Studio)","vendor":"Google","category":"speech-to-speech","summary":"Stateful WebSocket API for Gemini native-audio models with audio, text and 1 fps image/video input and spoken output. Much cheaper per token than OpenAI, with a free tier; suits prototypes and cost-sensitive voice agents that can live with short connection limits.","status":"GA","models":[{"name":"gemini-3.8-live","status":"GA","notes":"GA Sept 15, 2026. Default low-latency voice agent model; 131,072 input / 65,536 output tokens; async function calling by default; thinkingLevel not supported (omit it). Supports Search grounding; no caching, structured outputs or code execution."},{"name":"gemini-3.8-live-extended-thinking","status":"GA","notes":"Background reasoning (thinkingLevel low/medium/high); only NON_BLOCKING function calls; turnComplete does not mean idle, check interaction_status."},{"name":"gemini-3.1-flash-live-preview","status":"Preview","notes":"Released Mar 26, 2026; now legacy, Google recommends moving to 3.8 Live. No affective dialog or proactive audio; sequential function calling."},{"name":"gemini-2.5-flash-native-audio-preview-12-2025","status":"Preview","notes":"Not deprecated but since Sept 18, 2026 limited to projects that used it before."},{"name":"gemini-3.5-live-translate-preview","status":"Preview","notes":"Live speech translation."},{"name":"gemini-3.5-transcribe-live","status":"Preview","notes":"Streaming STT over the Live API (Aug 26, 2026)."},{"name":"gemini-2.0-flash-live-001, gemini-live-2.5-flash-preview","status":"Deprecated","notes":"Shut down December 9, 2025."}],"transports":["WebSocket"],"audio":{"input":"Raw 16-bit PCM little-endian, natively 16 kHz (other rates resampled if the MIME type says so, e.g. audio/pcm;rate=16000). Images/video as JPEG or PNG frames, max 1 fps.","output":"Raw 16-bit PCM little-endian at 24 kHz (always)."},"languages":"Live guide lists 99 languages (overview page says 70); native audio models pick the language automatically and do not accept a language code.","voices":"30 prebuilt voices shared with Gemini TTS (e.g. Kore, Puck, Zephyr).","latency":"No numeric vendor claim found.","features":["native audio output","barge-in (serverContent.interrupted)","automatic VAD with start/end sensitivity, prefix padding, silence duration; or manual activityStart/activityEnd","function calling (async/NON_BLOCKING on 3.8) and Google Search grounding","input and output audio transcription","affective dialog (v1beta, not on 3.1 Flash Live)","proactive audio (model may choose not to answer)","thinking on extended-thinking and 3.1 Flash Live models","image and video frame input (1 fps)","context window compression (sliding window)","session resumption handles and GoAway warnings","ephemeral tokens with config locking (liveConnectConstraints)"],"pricing":{"model":"per-token","items":[{"what":"gemini-3.8-live (and extended thinking, 3.1 Flash Live) text input","price":"$0.75","unit":"per 1M tokens","notes":"paid tier"},{"what":"audio input","price":"$3.00","unit":"per 1M tokens","notes":"Google also quotes $0.005/min"},{"what":"image / video input","price":"$1.00","unit":"per 1M tokens","notes":"or $0.002/min"},{"what":"text output (incl. thinking)","price":"$4.50","unit":"per 1M tokens","notes":""},{"what":"audio output","price":"$12.00","unit":"per 1M tokens","notes":"or $0.018/min"},{"what":"gemini-2.5-flash-native-audio-preview-12-2025","price":"$0.50 text in, $3.00 audio/video in, $2.00 text out, $12.00 audio out","unit":"per 1M tokens","notes":"legacy access only"},{"what":"gemini-3.5-live-translate-preview","price":"$3.50 in / $21.00 out","unit":"per 1M audio tokens","notes":"about $0.0053/min in and $0.0315/min out"},{"what":"gemini-3.5-transcribe-live","price":"$3.50 audio in / $21.00 text out","unit":"per 1M tokens","notes":"about $0.009/min blended"},{"what":"Google Search grounding (Gemini 3+)","price":"5,000 free requests/month, then $14","unit":"per 1,000 search queries","notes":"each search query billed"},{"what":"Free tier","price":"$0","unit":"","notes":"all Live models free of charge on the free tier, rate limited; content used to improve Google products"}],"audio_token_rate":"25 tokens per second of audio in and out = 1,500 tokens/min (Google pricing footnotes and Vertex Live billing details); 258 tokens per image.","est_per_minute_usd":{"low":0.0225,"high":0.12,"basis":"Low = gemini-3.8-live, 1 min user audio (1,500 tokens x $3/1M) + 1 min model audio (1,500 tokens x $12/1M), single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Live models do not support context caching, so there is no cached discount. Video input or thinking tokens add more."},"free_tier":"Yes: free tier for all Live models with lower rate limits; free-tier content may be used to improve Google products.","source":"https://ai.google.dev/gemini-api/docs/pricing"},"limits":["Connection lifetime around 10 minutes (GoAway with timeLeft is sent before close)","Without compression: audio-only sessions 15 minutes, audio+video 2 minutes","Context window 128k tokens for native audio models (3.8 Live lists 131,072 input)","Session resumption tokens valid 2 hours after the last session ends (Developer API)","Ephemeral tokens: new-session window default 1 minute, message window default 30 minutes","Rate limits are per project and shown in AI Studio; Live concurrency is not published on the rate-limit page"],"regions":"No region selection on the Developer API (Google-managed global serving). Use Vertex AI for regional control and CMEK.","setup":{"steps":["Sign in to Google AI Studio (aistudio.google.com) and create an API key for a Google Cloud project.","Start on the free tier; enable billing on the project to move to paid tiers and keep data out of product improvement.","Install the SDK (npm i @google/genai or pip install google-genai).","Server-to-server: connect with the API key. Client-to-server: your server calls client.authTokens.create (v1beta) with uses: 1 and liveConnectConstraints, and the browser uses token.name as its apiKey.","Enable contextWindowCompression and sessionResumption from day one and handle goAway by reconnecting with the latest handle.","Stream 16 kHz PCM, play 24 kHz PCM, and stop playback on serverContent.interrupted."],"endpoint":"wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent?key=API_KEY (ephemeral tokens use BidiGenerateContentConstrained with access_token)","auth":"API key on the server only. Browsers and mobile use ephemeral tokens from POST https://generativelanguage.googleapis.com/v1beta/auth_tokens, sent as access_token query parameter or Authorization: Token <token>; lock model and config with liveConnectConstraints.","snippet_lang":"javascript","snippet":"import { GoogleGenAI, Modality } from \"@google/genai\";\n// Server side with an API key. In a browser, create an ephemeral token on your server\n// (client.authTokens.create) and pass token.name as apiKey instead.\nconst ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });\nconst session = await ai.live.connect({\n  model: \"gemini-3.8-live\",\n  config: {\n    responseModalities: [Modality.AUDIO],\n    systemInstruction: \"You are a friendly support agent. Keep answers short.\",\n    speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: \"Kore\" } } },\n    contextWindowCompression: { slidingWindow: {} }, // lifts the 15 min audio session cap\n    sessionResumption: {},                            // survive the ~10 min connection limit\n    outputAudioTranscription: {},\n  },\n  callbacks: {\n    onmessage: (msg) => {\n      for (const part of msg.serverContent?.modelTurn?.parts ?? []) // a message can hold several parts\n        if (part.inlineData?.data) playPcm24k(Buffer.from(part.inlineData.data, \"base64\"));\n      if (msg.serverContent?.interrupted) stopPlayback(); // barge-in\n      if (msg.sessionResumptionUpdate?.newHandle) saveHandle(msg.sessionResumptionUpdate.newHandle);\n      if (msg.goAway) scheduleReconnect(msg.goAway.timeLeft);\n    },\n    onerror: (e) => console.error(e),\n    onclose: (e) => console.log(\"closed\", e.reason),\n  },\n});\n// 16 kHz mono PCM16 chunks from the microphone\nexport const sendAudio = (pcm16) =>\n  session.sendRealtimeInput({ audio: { data: pcm16.toString(\"base64\"), mimeType: \"audio/pcm;rate=16000\" } });"},"warnings":[{"severity":"high","title":"10 minute connections, 15 minute sessions","detail":"Connections drop around 10 minutes and audio-only sessions stop at 15 minutes (2 minutes with video) unless you enable context window compression and session resumption. Every production app needs reconnect logic driven by goAway."},{"severity":"high","title":"Whole context re-billed every turn, no caching","detail":"Google bills all tokens in the session context window on each turn and Live models do not support caching. At 25 tokens/s, long calls grow fast; use a sliding window with a sensible trigger_tokens."},{"severity":"high","title":"Free tier trains on your data","detail":"Free-tier content is used to improve Google products. Do not run real customer calls on a free-tier key."},{"severity":"medium","title":"Model churn","detail":"Google moved from 2.0 Flash Live and Live 2.5 Flash (shut down Dec 9, 2025) to 2.5 native audio, 3.1 Flash Live preview (Mar 2026) and 3.8 Live (Sept 2026). The 2.5 models are now restricted to previous users. Config differences (thinkingLevel rejected on 3.8, async tools default) break copy-pasted setups."},{"severity":"medium","title":"Audio-only responses","detail":"Native audio models only support the AUDIO response modality. For text you must enable output audio transcription, and a single server event can carry several parts, so loop over all parts."},{"severity":"medium","title":"Manual VAD is unforgiving","detail":"With automatic VAD off the server adds no pre-speech buffer or silence tolerance; keep at least 500 ms end-of-speech silence and send audioStreamEnd when the mic pauses for over a second."},{"severity":"medium","title":"Ordering vs responsiveness","detail":"send_realtime_input favours speed over strict ordering; use send_client_content when order matters. On 3.8 models turn_complete=true in send_client_content interrupts generation."},{"severity":"low","title":"No echo cancellation server side","detail":"Use browser echoCancellation or headphones; otherwise the model hears itself and barges in on its own speech."}],"best_for":"Low-cost voice agents, prototypes on the free tier, and multimodal assistants that also look at a camera or screen (1 fps).","open_source":false,"self_hostable":false,"compliance":"Paid tier data is not used to improve products; free tier is. No BAA or data residency on the Developer API; use Vertex AI for enterprise compliance (CMEK, VPC-SC, regional processing).","docs":[{"label":"Live API overview","url":"https://ai.google.dev/gemini-api/docs/live"},{"label":"Live capabilities guide","url":"https://ai.google.dev/gemini-api/docs/live-guide"},{"label":"Session management","url":"https://ai.google.dev/gemini-api/docs/live-session"},{"label":"Ephemeral tokens","url":"https://ai.google.dev/gemini-api/docs/ephemeral-tokens"},{"label":"WebSocket get started","url":"https://ai.google.dev/gemini-api/docs/live-api/get-started-websocket"},{"label":"Pricing","url":"https://ai.google.dev/gemini-api/docs/pricing"}],"sources":["https://ai.google.dev/gemini-api/docs/pricing","https://ai.google.dev/gemini-api/docs/models","https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live","https://ai.google.dev/gemini-api/docs/live","https://ai.google.dev/gemini-api/docs/live-guide","https://ai.google.dev/gemini-api/docs/live-session","https://ai.google.dev/gemini-api/docs/ephemeral-tokens","https://ai.google.dev/gemini-api/docs/live-api/get-started-websocket","https://ai.google.dev/gemini-api/docs/changelog","https://ai.google.dev/gemini-api/docs/speech-generation","https://cloud.google.com/vertex-ai/generative-ai/pricing"],"confidence":"high","unverified":"Live API concurrent session limits per tier on the Developer API; exact language count (99 vs 70 on different pages); whether 'half-cascade' models still exist (current docs only describe native audio); latency figures. The live guide still says 'the Live API is in preview' while the models are labelled stable/GA.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Gemini Live","facts":{"latency_ms":null,"languages":99,"max_session_min":10,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":false,"soc2":null,"gdpr_eu":false,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":true,"voices":30,"voice_cloning":null,"context_tokens":131072,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":3,"price_audio_out_per_1m":12,"price_per_min":null,"_notes":"Connections drop at about 10 min; audio-only sessions stop at 15 min (2 min with video) without compression and resumption. 99 languages per the Live guide (overview page says 70). Free tier data is used to improve Google products. No BAA or region choice on the Developer API."}},{"id":"google-vertex-live-api","name":"Gemini Live API on Vertex AI (Gemini Enterprise Agent Platform)","vendor":"Google Cloud","category":"speech-to-speech","summary":"The same Gemini Live models served from Google Cloud with IAM auth, regional endpoints, CMEK, provisioned throughput and live avatar video output. Suits enterprises that need Google Cloud compliance and higher concurrency.","status":"GA","models":[{"name":"gemini-3.8-live","status":"GA","notes":"Recommended. Native audio, transcriptions, VAD, affective dialog, proactive audio, tool use and live avatars."},{"name":"gemini-live-2.5-flash-native-audio","status":"GA","notes":"Previous generation; same features minus live avatars. Migration guide to 3.8 Live exists."},{"name":"gemini-3.5-transcribe-live-preview","status":"Preview","notes":"Speech-to-text only with interim/final transcripts and timestamps."}],"transports":["WebSocket"],"audio":{"input":"Raw 16-bit PCM 16 kHz little-endian; JPEG images/video at 1 fps; text.","output":"Raw 16-bit PCM 24 kHz little-endian; text; mp4 video for live avatars."},"languages":"Vertex overview states 24 languages for multilingual support (Developer API docs claim more); verify per language.","voices":"Same Gemini prebuilt voice set (30 voices on the TTS list); not separately documented for Vertex in the pages read.","latency":"No numeric vendor claim found.","features":["native audio, barge-in, VAD","affective dialog and proactive audio","function calling and Google Search grounding","audio transcriptions","live avatars (video output) on gemini-3.8-live","context window compression and session resumption","Provisioned Throughput for Live API","CMEK in us and eu multi-regions"],"pricing":{"model":"per-token","items":[{"what":"Gemini 3.8 Live text input","price":"$0.75","unit":"per 1M tokens","notes":"non-global"},{"what":"Gemini 3.8 Live video / image input","price":"$1.00","unit":"per 1M tokens","notes":"258 tokens per image"},{"what":"Gemini 3.8 Live audio input","price":"$3.00","unit":"per 1M tokens","notes":""},{"what":"Gemini 3.8 Live text output (response and reasoning)","price":"$4.50","unit":"per 1M tokens","notes":""},{"what":"Gemini 3.8 Live audio output","price":"$12.00","unit":"per 1M tokens","notes":""},{"what":"Gemini 3.8 Live avatar video output","price":"$1.00","unit":"per 1M tokens","notes":"6,192 tokens per second of video while the avatar speaks (about $0.37/min); idle listening not billed"},{"what":"Gemini 2.5 Flash Live API","price":"$0.50 text in, $3.00 audio in, $3.00 video/image in, $2.00 text out, $12.00 audio out","unit":"per 1M tokens","notes":""},{"what":"Audio-to-text transcription","price":"text output rate","unit":"per 1M tokens","notes":"charged for transcript tokens when transcription is enabled"}],"audio_token_rate":"Audio 25 tokens per second (input and output) = 1,500 tokens/min; image 258 tokens; avatar video 6,192 tokens per second (Vertex Live API billing details).","est_per_minute_usd":{"low":0.0225,"high":0.12,"basis":"Low = 1 min user audio + 1 min model audio at $3 / $12 per 1M, single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Google states tokens from past turns are re-processed and billed every turn up to the context window limit. Live avatar video adds about $0.37 per minute of avatar speech."},"free_tier":"No Live-specific free tier; standard Google Cloud new-customer credits apply.","source":"https://cloud.google.com/vertex-ai/generative-ai/pricing"},"limits":["Up to 1,000 concurrent sessions per project on pay-as-you-go (not applied to Provisioned Throughput)","Connection lifetime around 10 minutes","Audio-only sessions 15 minutes and audio-video 2 minutes without context window compression","Context window limit 128k tokens; compression trigger 5,000 to 128,000 tokens","Session resumption within 24 hours (state typically kept around 10 minutes after disconnect per the same page)","Global region is not supported for Live API CMEK (us and eu multi-regions only)"],"regions":"Regional WebSocket endpoints ({LOCATION}-aiplatform.googleapis.com); CMEK serving profiles in us and eu multi-regions. Check the model's location list before choosing a region.","setup":{"steps":["Create or pick a Google Cloud project with billing enabled and enable the Vertex AI API.","Grant your service account the Vertex AI User role; authenticate with Application Default Credentials.","Install google-genai and create the client with vertexai=True, project and a supported location.","Connect with client.aio.live.connect(model='gemini-3.8-live', config=...) and enable context window compression and session resumption.","For browsers, proxy audio through your backend (or an ADK / partner media server); do not ship service account credentials.","For higher, guaranteed capacity buy Provisioned Throughput for Live API."],"endpoint":"wss://{LOCATION}-aiplatform.googleapis.com/ws/google.cloud.aiplatform.v1.LlmBidiService/BidiGenerateContent","auth":"Google Cloud IAM OAuth access token (ADC / service account) on your backend. No browser-safe key; keep the WebSocket on the server or behind a relay.","snippet_lang":"python","snippet":"import asyncio\nfrom google import genai\nfrom google.genai import types\n\n# Auth: Application Default Credentials (gcloud auth application-default login or a service account)\nclient = genai.Client(vertexai=True, project=\"YOUR_PROJECT\", location=\"us-central1\")  # pick a region that lists the model\nconfig = types.LiveConnectConfig(\n    response_modalities=[\"AUDIO\"],\n    system_instruction=\"You are a friendly support agent. Keep answers short.\",\n    context_window_compression=types.ContextWindowCompressionConfig(sliding_window=types.SlidingWindow()),\n    session_resumption=types.SessionResumptionConfig(),\n)\n\nasync def run(mic_chunks, play, stop_playback):\n    async with client.aio.live.connect(model=\"gemini-3.8-live\", config=config) as session:\n        async def pump():\n            async for chunk in mic_chunks():  # raw 16 kHz mono PCM16 bytes\n                await session.send_realtime_input(\n                    audio=types.Blob(data=chunk, mime_type=\"audio/pcm;rate=16000\"))\n        asyncio.create_task(pump())\n        async for msg in session.receive():\n            if msg.data:\n                play(msg.data)  # 24 kHz mono PCM16\n            if msg.server_content and msg.server_content.interrupted:\n                stop_playback()  # user barged in\n            if msg.go_away:\n                print(\"connection ending in\", msg.go_away.time_left)  # reconnect with the resumption handle"},"warnings":[{"severity":"high","title":"Context re-billing is explicit","detail":"Vertex documents that every turn bills all tokens in the session context window, including all previous turns. Long sessions with resumption and compression still pay for whatever stays in the window."},{"severity":"high","title":"Session clocks","detail":"~10 minute connections and 15 minute audio-only sessions (2 minutes with video) unless compression is on. Implement GoAway handling and resumption before going live."},{"severity":"medium","title":"GA vs preview labels disagree","detail":"Model docs mark gemini-3.8-live GA while the pricing page still footnotes 'Live API is in Preview'. Confirm SLA coverage with Google before committing to uptime targets."},{"severity":"medium","title":"Avatar video is expensive","detail":"Avatar output is 6,192 tokens per second; at $1/1M that is about $0.37 per minute of speaking time, roughly 15x the audio cost of a turn."},{"severity":"medium","title":"Docs moved","detail":"Vertex generative AI docs now redirect into the Gemini Enterprise Agent Platform site and the pricing page is titled Agent Platform Pricing. Old bookmarks and blog links may 404."},{"severity":"low","title":"Language support is narrower than the Developer API docs suggest","detail":"Vertex lists 24 languages for Live; test before promising other languages."}],"best_for":"Enterprises on Google Cloud that need IAM, CMEK, VPC-SC, provisioned throughput, high concurrency (1,000 sessions per project) or avatar video.","open_source":false,"self_hostable":false,"compliance":"Google Cloud terms; CMEK in us/eu multi-regions; customer data not used for training under Google Cloud terms. HIPAA BAA available for covered Google Cloud services (confirm Live API coverage).","docs":[{"label":"Live API overview (Vertex)","url":"https://docs.cloud.google.com/vertex-ai/generative-ai/docs/live-api"},{"label":"Start and manage live sessions","url":"https://docs.cloud.google.com/vertex-ai/generative-ai/docs/live-api/start-manage-session"},{"label":"Pricing","url":"https://cloud.google.com/vertex-ai/generative-ai/pricing"}],"sources":["https://docs.cloud.google.com/vertex-ai/generative-ai/docs/live-api","https://docs.cloud.google.com/vertex-ai/generative-ai/docs/live-api/start-manage-session","https://cloud.google.com/vertex-ai/generative-ai/pricing"],"confidence":"medium","unverified":"Exact list of regions serving gemini-3.8-live (example code uses us-central1 as a placeholder); whether Vertex Live supports ephemeral/browser tokens; voice list on Vertex; per-region price differences (the 3.8 Live rows are labelled Non-global).","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Gemini Live on Vertex","facts":{"latency_ms":null,"languages":24,"max_session_min":10,"concurrency":1000,"free_tier":false,"free_credit_usd":null,"hipaa":true,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":true,"voices":30,"voice_cloning":null,"context_tokens":128000,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":3,"price_audio_out_per_1m":12,"price_per_min":null,"_notes":"Connections about 10 min; audio-only sessions 15 min without compression. 1,000 concurrent sessions per project on pay-as-you-go. Vertex lists 24 languages. HIPAA BAA for covered Google Cloud services; confirm Live API coverage. Voice count assumed shared with Gemini TTS list. Generic Google Cloud new-customer credits may apply."}},{"id":"amazon-nova-sonic","name":"Amazon Nova 2 Sonic / Nova 2.5 Sonic (Bedrock)","vendor":"Amazon Web Services","category":"speech-to-speech","summary":"Amazon's native speech-to-speech models on Bedrock, streamed over HTTP/2 bidirectional InvokeModelWithBidirectionalStream. Cheap per token, strong for AWS-native contact centers (Amazon Connect), but with an 8 minute connection limit and a low fixed concurrency quota.","status":"GA","models":[{"name":"amazon.nova-2-sonic-v1:0","status":"GA","notes":"GA Dec 2, 2025; refreshed in place Mar 2026 (p50 latency -150 ms, Polly-compatible voices, better 8 kHz turn-taking) and May 2026 (88 percent fewer speech hallucinations on AWS internal set). 1M token context, 64K max output. EOL no sooner than Dec 2, 2026."},{"name":"amazon.nova-2-5-sonic","status":"GA","notes":"GA Oct 5, 2026: better reasoning, instruction following and tool calling, lower latency, 256K context, same price as Nova 2 Sonic. Model id taken from the Strands Bidi Agents launch post; confirm in the Bedrock console."},{"name":"amazon.nova-sonic-v1:0","status":"GA","notes":"Original Nova Sonic (Apr 2025); 300K context, English-focused; slightly higher speech prices. Superseded."}],"transports":["HTTP/2 bidirectional (InvokeModelWithBidirectionalStream)"],"audio":{"input":"audio/lpcm 16-bit mono at 8, 16 or 24 kHz, base64 in audioInput events (~32 ms frames streamed in real time).","output":"audio/lpcm 16-bit mono at 8, 16 or 24 kHz, base64 audioOutput events."},"languages":"English (US, UK, India, Australia), French, Italian, German, Spanish, Portuguese, Hindi, with automatic language detection and switching.","voices":"16 voice ids in the docs (matthew, tiffany, amy, olivia, lupe, carlos, ambre, florian, lennart, beatrice, lorenzo, tina, carolina, leo, kiara, arjun); polyglot voices can speak every supported language; Polly-compatible voices since March 2026.","latency":"Vendor claim: March 2026 refresh cut user-perceived p50 latency by 150 ms; Nova 2.5 Sonic claimed lower latency than Nova 2 Sonic (no absolute numbers published).","features":["turn detection with endpointingSensitivity HIGH / MEDIUM / LOW","barge-in handling without losing context","function calling with asynchronous tool handling (agent keeps talking while tools run)","cross-modal input: text messages during a voice session","conversation history injection at session start","RAG via your own tools (Bedrock Knowledge Bases integration patterns)","telephony via Amazon Connect, Twilio, Vonage, AudioCodes; LiveKit and Pipecat integrations","Strands Bidi Agents (GA Oct 2026) handles reconnects across the 8 minute limit"],"pricing":{"model":"per-token","items":[{"what":"Nova 2 Sonic / Nova 2.5 Sonic speech input","price":"$0.003","unit":"per 1K tokens ($3.00 per 1M)","notes":"us-east-1"},{"what":"Nova 2 Sonic / Nova 2.5 Sonic speech output","price":"$0.012","unit":"per 1K tokens ($12.00 per 1M)","notes":"us-east-1"},{"what":"Nova 2 Sonic / Nova 2.5 Sonic text input","price":"$0.00033","unit":"per 1K tokens ($0.33 per 1M)","notes":"system prompt, history, tool results, cross-modal text"},{"what":"Nova 2 Sonic / Nova 2.5 Sonic text output","price":"$0.00275","unit":"per 1K tokens ($2.75 per 1M)","notes":"transcripts and tool calls"},{"what":"Nova Sonic (v1) speech input / output","price":"$0.0034 / $0.0136","unit":"per 1K tokens","notes":"text $0.00006 in / $0.00024 out per 1K"}],"audio_token_rate":"Not published by AWS. An AWS re:Post community answer says about 25 tokens per second (unverified). Measure with usageEvent speech token counts.","est_per_minute_usd":{"low":0.0225,"high":0.12,"basis":"UNVERIFIED token rate. Assuming ~25 speech tokens/s (community figure): low = 1 min user speech (1,500 x $3/1M) + 1 min model speech (1,500 x $12/1M) = $0.0225, single turn. High uses the same 10 minute, 50 turn model as other providers and assumes history is re-processed each turn like other S2S APIs (AWS does not document this). Third-party sites quote roughly $0.015/min."},"free_tier":"None specific to Nova Sonic (AWS Free Tier credits may apply to new accounts).","source":"https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonBedrock/current/us-east-1/index.json (AWS Price List API, retrieved 2026-10-10)"},"limits":["Connection limit 8 minutes; renew the connection and carry history forward (AWS samples and Strands Bidi Agents do this)","Default quota 20 concurrent InvokeModelWithBidirectionalStream sessions per account per Region for Nova 2 Sonic, documented as not adjustable","Context window: Nova 2 Sonic 1M tokens (64K output); Nova 2.5 Sonic 256K per launch post","Conversation history can only be injected once, after the system prompt and before audio streaming","In-Region inference only (no geo or global cross-region profiles)","Standard service tier only (no Priority, Flex or Reserved)"],"regions":"us-east-1 (N. Virginia), us-west-2 (Oregon), eu-north-1 (Stockholm), ap-northeast-1 (Tokyo). Through Amazon Connect also Singapore, London, Seoul, Frankfurt (Nova 2 Sonic).","setup":{"steps":["Create an AWS account and pick a supported Region (e.g. us-east-1).","In the Bedrock console confirm access to Amazon Nova 2 Sonic or Nova 2.5 Sonic (and accept any model terms).","Create IAM credentials with bedrock:InvokeModelWithBidirectionalStream, or a Bedrock API key (AWS_BEARER_TOKEN_BEDROCK) for quick tests.","Use an SDK with HTTP/2 bidirectional streaming support (AWS SDK for JavaScript v3 with NodeHttp2Handler, Python aws_sdk_bedrock_runtime, Java, .NET); boto3 does not stream bidirectionally.","Send sessionStart, promptStart, system prompt content, then a long-lived AUDIO content block of audioInput events; read audioOutput, textOutput, toolUse and usageEvent events.","Plan reconnects before 8 minutes and request no more than 20 concurrent sessions per Region per account (use more accounts/Regions or Amazon Connect for scale)."],"endpoint":"https://bedrock-runtime.{region}.amazonaws.com (InvokeModelWithBidirectionalStream over HTTP/2), modelId amazon.nova-2-sonic-v1:0 or amazon.nova-2-5-sonic","auth":"AWS SigV4 with IAM credentials (or a Bedrock API key). There is no browser token flow: run the stream on your backend and relay audio from the browser or phone over WebSocket/WebRTC.","snippet_lang":"javascript","snippet":"import { BedrockRuntimeClient, InvokeModelWithBidirectionalStreamCommand } from \"@aws-sdk/client-bedrock-runtime\";\nimport { NodeHttp2Handler } from \"@smithy/node-http-handler\";\nimport { randomUUID } from \"node:crypto\";\n\nconst client = new BedrockRuntimeClient({ region: \"us-east-1\", requestHandler: new NodeHttp2Handler({ requestTimeout: 300000 }) });\nconst P = randomUUID(), SYS = randomUUID(), MIC = randomUUID();\nconst enc = (event) => ({ chunk: { bytes: new TextEncoder().encode(JSON.stringify({ event })) } });\nconst audioCfg = { mediaType: \"audio/lpcm\", sampleRateHertz: 16000, sampleSizeBits: 16, channelCount: 1, audioType: \"SPEECH\", encoding: \"base64\" };\n\nasync function* input(micChunks) {\n  yield enc({ sessionStart: { inferenceConfiguration: { maxTokens: 1024, topP: 0.9, temperature: 0.7 }, turnDetectionConfiguration: { endpointingSensitivity: \"MEDIUM\" } } });\n  yield enc({ promptStart: { promptName: P, textOutputConfiguration: { mediaType: \"text/plain\" },\n    audioOutputConfiguration: { ...audioCfg, sampleRateHertz: 24000, voiceId: \"matthew\" } } });\n  yield enc({ contentStart: { promptName: P, contentName: SYS, type: \"TEXT\", interactive: false, role: \"SYSTEM\", textInputConfiguration: { mediaType: \"text/plain\" } } });\n  yield enc({ textInput: { promptName: P, contentName: SYS, content: \"You are a friendly support agent. Keep answers short.\" } });\n  yield enc({ contentEnd: { promptName: P, contentName: SYS } });\n  yield enc({ contentStart: { promptName: P, contentName: MIC, type: \"AUDIO\", interactive: true, role: \"USER\", audioInputConfiguration: audioCfg } });\n  for await (const pcm of micChunks()) // 16 kHz mono PCM16, ~32 ms frames, streamed in real time\n    yield enc({ audioInput: { promptName: P, contentName: MIC, content: Buffer.from(pcm).toString(\"base64\") } });\n  yield enc({ contentEnd: { promptName: P, contentName: MIC } });\n  yield enc({ promptEnd: { promptName: P } });\n  yield enc({ sessionEnd: {} });\n}\n\nexport async function talk(micChunks, play) {\n  const res = await client.send(new InvokeModelWithBidirectionalStreamCommand({ modelId: \"amazon.nova-2-sonic-v1:0\", body: input(micChunks) }));\n  for await (const ev of res.body) {\n    if (!ev.chunk?.bytes) continue;\n    const msg = JSON.parse(new TextDecoder().decode(ev.chunk.bytes)).event;\n    if (msg?.audioOutput) play(Buffer.from(msg.audioOutput.content, \"base64\")); // 24 kHz PCM16\n  }\n}"},"warnings":[{"severity":"high","title":"20 concurrent sessions, not adjustable","detail":"The Nova 2 Sonic model card lists 20 concurrent bidirectional sessions per account per Region as a non-adjustable quota. That caps a single-account deployment at 20 simultaneous calls per Region; plan multi-Region, multi-account or Amazon Connect for real call volume."},{"severity":"high","title":"8 minute connection limit","detail":"Each bidirectional stream ends at 8 minutes. You must open a new stream and re-send system prompt and history (paying text input again) at a natural pause, or use Strands Bidi Agents which automates this."},{"severity":"medium","title":"Token rate not published","detail":"AWS bills speech tokens but does not publish tokens per second, so per-minute cost can only be measured from usageEvent data. Budget conservatively until you have pilot numbers."},{"severity":"medium","title":"Strict event ordering","detail":"Every event must carry the right promptName and contentName, history must come before audio, and sessions must close with contentEnd, promptEnd, sessionEnd. Mistakes cause opaque stream errors or orphaned sessions."},{"severity":"medium","title":"SDK support is uneven","detail":"Bidirectional streaming needs HTTP/2 SDK support; boto3 does not support it, and the Python path uses a separate experimental SDK. Keep audio frames flowing in real time or the stream can time out."},{"severity":"medium","title":"Silent in-place model refreshes","detail":"Nova 2 Sonic was updated in place in March and May 2026 with no API change. Behaviour (voices, turn-taking) can shift under a fixed model id; keep regression tests for prompts and tools."},{"severity":"low","title":"Limited languages","detail":"Seven languages; no Japanese, Chinese or Arabic despite the Tokyo Region. Check before selling into those markets."},{"severity":"low","title":"Bedrock features missing","detail":"Guardrails, Knowledge Bases, Agents, prompt management and token counting are not supported on the Nova 2 Sonic bidirectional endpoint; implement safety and retrieval in your tools."}],"best_for":"AWS-native contact centers and phone agents (Amazon Connect, Twilio, Vonage) that want low per-token cost and in-Region processing.","open_source":false,"self_hostable":false,"compliance":"Bedrock data is not used to train models and stays in the chosen Region (in-Region inference only). Bedrock is HIPAA eligible and covered by AWS SOC reports; confirm Nova Sonic is listed in your BAA scope.","docs":[{"label":"Nova 2 Sonic user guide","url":"https://docs.aws.amazon.com/nova/latest/nova2-userguide/using-conversational-speech.html"},{"label":"Input events reference","url":"https://docs.aws.amazon.com/nova/latest/nova2-userguide/sonic-input-events.html"},{"label":"Bedrock model card","url":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html"},{"label":"Release notes","url":"https://docs.aws.amazon.com/nova/latest/nova2-userguide/release-notes.html"},{"label":"Samples","url":"https://github.com/aws-samples/amazon-nova-samples/tree/main/speech-to-speech"},{"label":"Bedrock pricing","url":"https://aws.amazon.com/bedrock/pricing/"}],"sources":["https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonBedrock/current/us-east-1/index.json","https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html","https://docs.aws.amazon.com/nova/latest/nova2-userguide/using-conversational-speech.html","https://docs.aws.amazon.com/nova/latest/nova2-userguide/sonic-input-events.html","https://docs.aws.amazon.com/nova/latest/nova2-userguide/release-notes.md","https://aws.amazon.com/about-aws/whats-new/2026/10/amazon-nova-2.5-sonic/","https://aws.amazon.com/blogs/aws/introducing-amazon-nova-2-sonic-next-generation-speech-to-speech-model-for-conversational-ai/","https://strandsagents.com/blog/bidi-agents-now-ga/","https://repost.aws/questions/QUNhHJ1dstS9mzzE6ZyRepjg/how-long-a-sonic-token-is"],"confidence":"medium","unverified":"Speech tokens per second (only a community figure of ~25/s); whether history is re-billed per turn; official Nova 2.5 Sonic model id (amazon.nova-2-5-sonic from the Strands blog), its concurrency quota and model card; whether the 20-session quota also applies to Nova 2.5 Sonic; prices in Regions other than us-east-1.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Amazon Nova Sonic","facts":{"latency_ms":null,"languages":7,"max_session_min":8,"concurrency":20,"free_tier":false,"free_credit_usd":null,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":false,"websocket":false,"sip":false,"open_weights":false,"function_calling":true,"image_input":null,"voices":16,"voice_cloning":null,"context_tokens":1000000,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":3,"price_audio_out_per_1m":12,"price_per_min":null,"_notes":"Transport is HTTP/2 bidirectional streaming. 8 min connection limit. 20 concurrent sessions per account per Region, not adjustable (Nova 2 Sonic). Context 1M for Nova 2 Sonic, 256K for Nova 2.5 Sonic. HIPAA and SOC are Bedrock-level; confirm Nova Sonic in BAA scope. EU via eu-north-1. Telephony via Amazon Connect and partners."}},{"id":"xai-grok-voice-agent","name":"xAI Grok Voice Agent API","vendor":"xAI","category":"speech-to-speech","summary":"Realtime speech-to-speech WebSocket API for Grok voice models with built-in web search, X search, file search and MCP tools, billed at a flat per-minute rate. Largely OpenAI Realtime compatible, so it is an easy second provider.","status":"GA","models":[{"name":"grok-voice-think-fast-2.0","status":"GA","notes":"Current speech-to-speech model; alias grok-voice-latest points to it. Regions us-east-1, eu-west-1, us-saltlake-2."},{"name":"grok-voice-think-fast-1.0","status":"GA","notes":"Previous model ($0.05/min per third-party reports); must be pinned explicitly since the alias moved to 2.0 (reported Aug 5, 2026)."}],"transports":["WebSocket"],"audio":{"input":"audio/pcm at 8000, 16000, 22050, 24000 (default), 32000, 44100 or 48000 Hz; audio/pcmu, audio/pcma (G.711); audio/opus. JSON base64 or binary transport.","output":"Same format options as input; speed 0.7 to 1.5."},"languages":"Docs say every voice can speak every supported language; TTS lists 20 languages and STT 38+. No explicit list for the voice agent.","voices":"Built-in voices including eve (default), ara and rex (full list via GET /v1/tts/voices); custom voices cloned from a reference clip up to 120 s.","latency":"Vendor claim: sub-second latency.","features":["server_vad turn detection with optional idle_timeout_ms re-engagement","reasoning.effort high or none","tools: web_search, x_search, file_search (Collections), remote MCP, custom functions","custom cloned voices","session resumption (opt-in conversation caching)","replace (spoken substitutions before TTS), keyterms and language hints","binary audio transport","OpenAI Realtime client compatibility with a base URL change","partner integrations: LiveKit, Pipecat, Twilio, Voximplant"],"pricing":{"model":"per-minute","items":[{"what":"Speech to speech (grok-voice-think-fast-2.0)","price":"$0.08","unit":"per minute ($4.80 per hour)","notes":""},{"what":"Text input during a voice session","price":"$0.004","unit":"per text input","notes":"third-party sources describe it as a flat fee per conversation.item.create event"},{"what":"web_search tool","price":"$5","unit":"per 1,000 calls","notes":"xAI tools price list; application to voice sessions not stated"},{"what":"x_search tool","price":"$5 per 1,000 posts, $10 per 1,000 profiles","unit":"","notes":""},{"what":"Speech to text (separate API)","price":"$0.10 REST / $0.20 streaming","unit":"per hour","notes":""},{"what":"Text to speech (separate API)","price":"$15.00","unit":"per 1M characters","notes":""}],"audio_token_rate":"Not applicable (billed per minute).","est_per_minute_usd":{"low":0.08,"high":0.08,"basis":"Flat $0.08 per minute on grok-voice-think-fast-2.0 regardless of context length. Tool calls and text inputs are extra. xAI does not state whether minutes are session wall-clock time or audio time, so assume wall-clock time including silence."},"free_tier":"None documented for the voice agent.","source":"https://docs.x.ai/developers/models"},"limits":["Session duration and concurrency limits are not published","Resumption history dropped after 30 minutes of inactivity","keyterms: max 100 terms, 50 characters each","Ephemeral client secrets: expiry set by expires_after.seconds (example 300); session and anchor fields not supported"],"regions":"us-east-1, eu-west-1, us-saltlake-2; EU data residency options per xAI docs.","setup":{"steps":["Create an xAI account at console.x.ai, add billing credit and create an API key.","Server: connect to the realtime WebSocket with Authorization: Bearer <key> and send session.update.","Browser: your server calls POST https://api.x.ai/v1/realtime/client_secrets with {\"expires_after\": {\"seconds\": 300}} and returns the token; the browser opens the WebSocket with subprotocol xai-client-secret.<token>.","Stream input_audio_buffer.append and play response.output_audio.delta.","Pin grok-voice-think-fast-2.0 (or 1.0) instead of grok-voice-latest in production."],"endpoint":"wss://api.x.ai/v1/realtime?model=grok-voice-latest","auth":"API key as Authorization: Bearer on the server. Browsers use short-lived tokens from POST https://api.x.ai/v1/realtime/client_secrets passed in the Sec-WebSocket-Protocol header with the xai-client-secret. prefix.","snippet_lang":"javascript","snippet":"import WebSocket from \"ws\"; // server side; browsers use an ephemeral token via the sec-websocket-protocol header\nconst ws = new WebSocket(\"wss://api.x.ai/v1/realtime?model=grok-voice-latest\", {\n  headers: { Authorization: `Bearer ${process.env.XAI_API_KEY}` },\n});\nws.on(\"open\", () => ws.send(JSON.stringify({\n  type: \"session.update\",\n  session: {\n    voice: \"eve\",\n    instructions: \"You are a friendly support agent. Keep answers short.\",\n    turn_detection: { type: \"server_vad\" },\n    audio: {\n      input: { format: { type: \"audio/pcm\", rate: 24000 } },\n      output: { format: { type: \"audio/pcm\", rate: 24000 } },\n    },\n    tools: [{ type: \"web_search\" }],\n  },\n})));\nexport function sendAudio(pcm16) { // 24 kHz mono PCM16\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: pcm16.toString(\"base64\") }));\n}\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"response.output_audio.delta\" || ev.type === \"response.audio.delta\")\n    playPcm16(Buffer.from(ev.delta, \"base64\"));\n  if (ev.type === \"input_audio_buffer.speech_started\") stopPlayback();\n  if (ev.type === \"error\") console.error(ev.error);\n});"},"warnings":[{"severity":"high","title":"Alias moved and price rose","detail":"grok-voice-latest now routes to think-fast-2.0 at $0.08/min, up from $0.05/min on 1.0 (per third-party reports). Pin the model id you tested and priced."},{"severity":"medium","title":"Per-minute billing includes idle time","detail":"Flat per-minute billing means open but silent sessions likely cost money. Close sockets on hang-up and use idle timeouts."},{"severity":"medium","title":"Not fully OpenAI compatible","detail":"Transcription delta is conversation.item.input_audio_transcription.updated with a cumulative transcript; conversation.item.retrieve, rate_limits.updated and conversation.item.done are unsupported. Test your OpenAI client code paths."},{"severity":"medium","title":"Tool fees add up","detail":"Built-in web_search and x_search are convenient but priced per call or per result on the general tools price list. Cap tool usage in instructions and log it."},{"severity":"medium","title":"Thin published limits","detail":"No public session length, concurrency or rate-limit numbers. Load test and ask xAI sales for written limits before a launch."},{"severity":"low","title":"Default reasoning is high","detail":"reasoning.effort defaults to high; set none for snappier small talk if latency matters."}],"best_for":"Teams wanting simple flat per-minute pricing, built-in live web and X search, voice cloning, or an OpenAI-compatible fallback provider.","open_source":false,"self_hostable":false,"compliance":"xAI states SOC 2 Type II, HIPAA eligible with a BAA, GDPR with EU data residency options, and that audio is never stored or used for training (vendor claims).","docs":[{"label":"Voice overview","url":"https://docs.x.ai/docs/guides/voice"},{"label":"Voice agent API","url":"https://docs.x.ai/developers/model-capabilities/audio/voice-agent"},{"label":"Ephemeral tokens","url":"https://docs.x.ai/developers/model-capabilities/audio/ephemeral-tokens"},{"label":"Pricing","url":"https://docs.x.ai/developers/models"}],"sources":["https://docs.x.ai/docs/guides/voice","https://docs.x.ai/developers/model-capabilities/audio/voice-agent","https://docs.x.ai/developers/model-capabilities/audio/ephemeral-tokens","https://docs.x.ai/developers/models","https://docs.x.ai/developers/pricing","https://docs.x.ai/developers/models/grok-voice-think-fast-2.0","https://www.eesel.ai/blog/grok-voice-think-fast-2-pricing"],"confidence":"medium","unverified":"How billable minutes are counted; whether tool fees apply inside voice sessions; session length and concurrency limits; 1.0 pricing and alias switch date (third-party); full voice and language lists.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Grok Voice Agent","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":true,"context_tokens":null,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.08,"_notes":"Vendor claims sub-second latency (no ms figure). No language list for the voice agent (TTS 20, STT 38+). Compliance items are vendor claims. Tool calls and text inputs billed extra. Telephony via partners (Twilio, Voximplant)."}},{"id":"alibaba-qwen-omni-realtime","name":"Alibaba Cloud Model Studio Qwen-Omni-Realtime","vendor":"Alibaba Cloud","category":"speech-to-speech","summary":"Realtime audio and video conversation API for Qwen Omni models over WebSocket or WebRTC, with an OpenAI-style event protocol. By far the cheapest frontier option per minute and strong in Chinese and Asian languages.","status":"GA","models":[{"name":"qwen3.8-omni-flash-realtime","status":"GA","notes":"Recommended; WebSocket, WebRTC and AOQ; 196,608 input tokens; workspace id required."},{"name":"qwen3.5-omni-plus-realtime","status":"GA","notes":"Larger, much more expensive model."},{"name":"qwen3.5-omni-flash-realtime","status":"GA","notes":"Previous flash model, shorter retained history."},{"name":"qwen3-omni-flash-realtime(-2025-09-15)","status":"GA","notes":"Older; 12.5 tokens/s for both input and output audio."}],"transports":["WebSocket","WebRTC"],"audio":{"input":"PCM 16 kHz (16-bit mono LE); multichannel 2 or 4 channel spatial input supported (double tokens); JPG images about 1 fps for video.","output":"PCM 24 kHz."},"languages":"Speech recognition 113 languages and dialects; speech generation 36 languages and dialects.","voices":"Multiple built-in voices (default Tina; others include Ethan, longanlingxin) plus voice cloning.","latency":"No numeric vendor claim captured.","features":["server_vad, semantic_vad (filters backchannels and noise) or manual turns (WebSocket only)","function calling and remote MCP tools (no extra MCP fee)","web search (cannot be combined with tool calling)","image and video frame input","voice cloning","multichannel spatial audio input"],"pricing":{"model":"per-token","items":[{"what":"qwen3.8-omni-flash-realtime text/image/video input","price":"$0.23","unit":"per 1M tokens","notes":"Singapore (International)"},{"what":"qwen3.8-omni-flash-realtime audio input","price":"$0.93","unit":"per 1M tokens","notes":""},{"what":"qwen3.8-omni-flash-realtime text output","price":"$0.70","unit":"per 1M tokens","notes":"speech output bills audio AND its matching text"},{"what":"qwen3.8-omni-flash-realtime audio output","price":"$1.87","unit":"per 1M tokens","notes":""},{"what":"qwen3.5-omni-plus-realtime","price":"$2.10 text in, $16.50 audio in, $12.40 text out, $62.00 audio out","unit":"per 1M tokens","notes":"only audio billed for speech output"},{"what":"qwen3.5-omni-flash-realtime","price":"$0.55 text in, $4.50 audio in, $3.30 text out, $17.70 audio out","unit":"per 1M tokens","notes":""},{"what":"Free quota","price":"1M tokens per model","unit":"","notes":"Singapore only, valid 90 days from activation or model release"}],"audio_token_rate":"qwen3.8-omni-flash-realtime and Qwen3.5-Omni-Realtime: input 7 tokens/s (420/min), output 12.5 tokens/s (750/min); qwen3-omni-flash-realtime-2025-09-15: 12.5 tokens/s both ways. Audio under 1 s billed as 1 s.","est_per_minute_usd":{"low":0.0018,"high":0.015,"basis":"Low = qwen3.8-omni-flash-realtime, 1 min user audio (420 tokens x $0.93/1M) + 1 min model audio (750 tokens x $1.87/1M), single turn, excluding the matching output text tokens (small). High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Alibaba documents that each turn re-bills retained history. qwen3.5-omni-plus-realtime: $0.053 low, $0.27 high."},"free_tier":"1M free tokens per model in the Singapore region for 90 days.","source":"https://www.alibabacloud.com/help/en/model-studio/model-pricing"},"limits":["Single WebSocket session up to 120 minutes","qwen3.8-omni-flash-realtime: up to 196,608 input tokens; retains 100 audio turns, 50 video turns, 600 s of audio and 240 s of video","qwen3.5-omni-flash-realtime retains 80 audio turns, 480 s audio, 120 s video","Base64 images under 256 KB; max 1080p","Concurrency limits on a separate rate-limiting page (not captured)"],"regions":"Singapore (ap-southeast-1, International) and China Beijing (cn-beijing); separate API keys per region.","setup":{"steps":["Create an Alibaba Cloud account, activate Model Studio in the Singapore (International) region, and note your workspace id.","Create a Model Studio (DashScope) API key for that region.","Server: open the WebSocket with Authorization: Bearer <key> and ?model=qwen3.8-omni-flash-realtime.","Send session.update (modalities, voice, instructions, turn_detection), stream input_audio_buffer.append, play response.audio.delta.","Browser: use the WebRTC SDP exchange at /api/v1/webrtc/realtime via your server (VAD mode only)."],"endpoint":"wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime (Beijing: cn-beijing host); WebRTC: https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/webrtc/realtime","auth":"Bearer API key on your server; each region needs its own key. No documented browser ephemeral token: relay or proxy the WebRTC SDP exchange through your backend.","snippet_lang":"javascript","snippet":"import WebSocket from \"ws\"; // server side; keep the DashScope / Model Studio key off the client\nconst WS_ID = process.env.MODEL_STUDIO_WORKSPACE_ID; // required for qwen3.8-omni-flash-realtime\nconst url = `wss://${WS_ID}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime`;\nconst ws = new WebSocket(url, { headers: { Authorization: `Bearer ${process.env.DASHSCOPE_API_KEY}` } });\nws.on(\"open\", () => ws.send(JSON.stringify({\n  type: \"session.update\",\n  session: {\n    modalities: [\"text\", \"audio\"],\n    voice: \"Tina\",\n    instructions: \"You are a friendly support agent. Keep answers short.\",\n    turn_detection: { type: \"semantic_vad\" },\n  },\n})));\nexport function sendAudio(pcm16) { // input must be 16 kHz mono PCM16\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: pcm16.toString(\"base64\") }));\n}\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"response.audio.delta\") playPcm24k(Buffer.from(ev.delta, \"base64\")); // 24 kHz PCM out\n  if (ev.type === \"input_audio_buffer.speech_started\") stopPlayback();\n  if (ev.type === \"error\") console.error(ev.error);\n});"},"warnings":[{"severity":"medium","title":"History re-billed per turn","detail":"Alibaba explicitly bills retained history again on each turn, so per-minute cost rises over a call; still very low on the 3.8 flash model."},{"severity":"medium","title":"Plus model is not cheap","detail":"qwen3.5-omni-plus-realtime audio output is $62/1M, about the same as OpenAI gpt-realtime-2.1. Do not assume all Qwen realtime models are budget options."},{"severity":"medium","title":"Regions and keys are separate","detail":"Singapore and Beijing use different hosts, keys and prices; the free quota is Singapore only. Mainland China endpoints carry their own data and regulatory implications."},{"severity":"medium","title":"Protocol differs between model generations","detail":"qwen3.8 examples use a nested session config while Qwen3.5 examples use flat input_audio_format/output_audio_format fields; output is response.audio.delta (beta-style naming), not OpenAI GA names."},{"severity":"low","title":"Feature conflicts","detail":"Web search cannot be combined with tool calling; WebRTC supports only VAD mode; representation_compact cannot change after audio starts."},{"severity":"high","title":"Context replay drives the bill","detail":"Input tokens accumulate: every response re-bills the retained prior audio, images and text as input. Long sessions on 3.5-plus can cost 10x the first minute. Keep sessions short or trim history."},{"severity":"high","title":"Model ids churn fast","detail":"In about a year the recommended id moved from qwen-omni-turbo-realtime to qwen3-omni-flash-realtime to qwen3.5 to qwen3.8. Older ids vanish from docs without a clear retirement notice on the page. Pin dated snapshots where offered and watch the Model Studio deprecation page."},{"severity":"medium","title":"3.8 bills audio and text for speech output","detail":"For qwen3.8-omni-flash-realtime, spoken output is billed as both audio tokens and the matching text tokens; 3.5 models bill only the audio. Spatial (multichannel) input doubles input-audio tokens."},{"severity":"medium","title":"Workspace endpoint and new SDK required","detail":"qwen3.8 only works on the workspace-specific host and needs DashScope Python SDK 1.26.5+ or Java 2.22.15+. Old dashscope-intl.aliyuncs.com examples on the web will not work for it."},{"severity":"medium","title":"Region choice is a compliance decision","detail":"Singapore and Beijing are separate deployments with separate keys and different prices. Beijing (Chinese mainland) means data processed in China and typically a China-verified account."},{"severity":"low","title":"Tools vs web search","detail":"Web search and tool calling are mutually exclusive in a session; MCP tools need a public HTTPS Streamable HTTP server and approval is on by default."},{"severity":"low","title":"Python SDK sends transcription by default","detail":"The Python SDK defaults enable_input_audio_transcription to True; set it explicitly if you do not want transcription events (and any associated cost)."},{"severity":"low","title":"Session hard cap","detail":"Sessions end at 120 minutes; build reconnect logic for long-running calls."}],"best_for":"Very cost-sensitive voice agents, Chinese and Asian-language markets, and audio-plus-video assistants.","open_source":false,"self_hostable":false,"compliance":"Data retention, residency guarantees and certifications for Model Studio realtime were not captured; review Alibaba Cloud International terms before sending regulated data.","docs":[{"label":"Qwen-Omni-Realtime guide","url":"https://www.alibabacloud.com/help/en/model-studio/realtime"},{"label":"Model pricing","url":"https://www.alibabacloud.com/help/en/model-studio/model-pricing"}],"sources":["https://www.alibabacloud.com/help/en/model-studio/realtime","https://www.alibabacloud.com/help/en/model-studio/model-pricing"],"confidence":"medium","unverified":"Concurrency limits, latency, data retention and compliance; full voice list; exact session.update shape for qwen3.8 (examples differ between generations); Beijing prices.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Qwen-Omni Realtime","facts":{"latency_ms":null,"languages":36,"max_session_min":120,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":false,"webrtc":true,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":true,"voices":null,"voice_cloning":true,"context_tokens":196608,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":0.93,"price_audio_out_per_1m":1.87,"price_per_min":null,"_notes":"36 speech-output languages; recognition covers 113. Prices are qwen3.8-omni-flash-realtime, Singapore. Free quota is 1M tokens per model for 90 days (Singapore only). Regions are Singapore and Beijing only."}},{"id":"volcengine-doubao-realtime","name":"Doubao Realtime Voice Model (end-to-end realtime speech)","vendor":"ByteDance Volcengine (Doubao Speech)","category":"speech-to-speech","summary":"ByteDance's end-to-end speech-to-speech model behind the Doubao app's voice chat, sold through Volcengine in mainland China. Version 3.0 'Seeduplex' is full-duplex with an OpenAI-Realtime-style JSON protocol; it suits Chinese-language consumer and companion products.","status":"GA","models":[{"name":"Doubao Realtime 3.0 (Seeduplex), session.model = 1.2.6.1","status":"GA","notes":"Full-duplex version, JSON text frames, function calling, endpoint /api/v3/duplex/realtime/dialogue."},{"name":"O2.0 / SC2.0 (half-duplex S2S)","status":"GA","notes":"Listed under the 'historical' end-to-end interface docs with a binary protocol. O = Omni route, SC = Strong Character role-play route; 12K max context."},{"name":"O / SC (1.x)","status":"GA","notes":"No longer iterated; capabilities are converging into the 2.0 versions."}],"transports":["WebSocket"],"audio":{"input":"PCM 16 kHz mono int16 little-endian (Opus also accepted and converted server-side); send 20 ms / 640-byte packets at real-time pace","output":"Ogg Opus by default; PCM 24 kHz mono (32-bit float or s16le) on request"},"languages":"Chinese and English (vendor says other languages are not guaranteed, especially for cloned voices).","voices":"O/O2.0: vv, xiaohe, yunzhou, xiaotian Chinese voices; O2.0 adds English voices Tim, Dacey, Stokie. SC/SC2.0: 21 official cloned character voices; custom voice cloning sold separately.","latency":"Vendor describes it as low latency; no millisecond figure on the API page.","features":["full-duplex (3.0)","barge-in","server VAD, push-to-talk, text input and audio-file input modes","function calling with parallel calls (3.0)","system prompt / persona fields","voice cloning (SC2.0, cloning 2.0 product)","singing","web search via extension","hot words","conversation history injection and resume (keeps last 20 rounds)"],"pricing":{"model":"per-token","items":[{"what":"Input audio","price":"80 CNY","unit":"per 1M tokens","notes":"Rates taken from the official worked billing example on the pricing page."},{"what":"Input text","price":"10 CNY","unit":"per 1M tokens","notes":""},{"what":"Cached input (text or audio)","price":"5 CNY","unit":"per 1M tokens","notes":"Context and system prompt hits from earlier turns."},{"what":"Output audio","price":"300 CNY","unit":"per 1M tokens","notes":""},{"what":"Output text","price":"30 CNY","unit":"per 1M tokens","notes":"Includes text of the spoken reply and returned ASR text."}],"audio_token_rate":"Input audio about 6.25 tokens per second; output audio about 25 tokens per second; ratios may change as the model updates (source: Doubao Speech billing page).","est_per_minute_usd":{"low":0.03,"high":0.07,"basis":"Own estimate at about 7.1 CNY per USD: 30 s user + 30 s agent speech is about 0.015 CNY input + 0.225 CNY output audio plus cached context, about 0.25-0.35 CNY/min; an agent speaking the full minute is about 0.45 CNY of output audio."},"free_tier":"Free quota exists and can offset cached, uncached and output tokens, but the size was not found in the docs read.","source":"https://docs.volcengine.com/docs/6561/1359370?lang=zh"},"limits":["Default QPM 60 (StartSession / session.create per minute per AppID) and TPM 100,000; raise via sales","Server releases the connection after 10 minutes with no interaction (error 45000003)","Uplink audio must keep real-time pace; send input_audio_mute.commit when the mic is muted or the session times out","Close with session.close and wait for the reply, otherwise error 55000001 ContextCanceled","O2.0/SC2.0 max context 12K"],"regions":"Chinese mainland (openspeech.bytedance.com). No equivalent speech-to-speech API was found on BytePlus (international).","setup":{"steps":["Create a Volcengine account and complete real-name verification (Chinese ID or business licence expected).","In the Doubao Speech console, enable the end-to-end realtime voice model and create an API key (new console).","Read the 'access must-read' page: 3.0 uses a new JSON event protocol that differs from the older binary protocol.","Open the WebSocket with X-Api-Key, send session.create with session.model = 1.2.6.1, then stream 20 ms PCM frames.","Handle response.output_audio.delta (Ogg Opus by default) and close with session.close."],"endpoint":"wss://openspeech.bytedance.com/api/v3/duplex/realtime/dialogue (3.0 full duplex); legacy: wss://openspeech.bytedance.com/api/v3/realtime/dialogue","auth":"3.0: X-Api-Key header (new console). Legacy: X-Api-App-ID, X-Api-Access-Key, X-Api-Resource-Id: volc.speech.dialog, X-Api-App-Key: PlgvMymc7f3tQnJ6 (fixed value)","snippet_lang":"python","snippet":"# pip install websockets   (Doubao Realtime 3.0 \"Seeduplex\", full-duplex JSON protocol)\nimport asyncio, base64, json, os, websockets\n\nURL = \"wss://openspeech.bytedance.com/api/v3/duplex/realtime/dialogue\"\nHEADERS = {\"X-Api-Key\": os.environ[\"VOLC_SPEECH_API_KEY\"]}  # new console > API Key management\n\nasync def main(pcm_chunks):  # 16 kHz mono int16 LE, 20 ms (640-byte) chunks, sent in real time\n    async with websockets.connect(URL, additional_headers=HEADERS) as ws:\n        await ws.send(json.dumps({\n            \"type\": \"session.create\",\n            \"session\": {\"model\": \"1.2.6.1\",  # fixed value for the full-duplex version\n                        \"instructions\": \"You are a friendly assistant.\"},\n        }))\n        async def uplink():\n            for chunk in pcm_chunks:\n                await ws.send(json.dumps({\"type\": \"input_audio_buffer.append\",\n                                          \"audio\": base64.b64encode(chunk).decode()}))\n                await asyncio.sleep(0.02)  # real-time pace: too fast or too slow is an error\n        async def downlink():\n            async for msg in ws:\n                ev = json.loads(msg)\n                if ev[\"type\"] == \"response.output_audio.delta\":\n                    pass  # base64 Ogg Opus by default; decode and play\n                elif ev[\"type\"] == \"error\":\n                    print(ev)\n        await asyncio.gather(uplink(), downlink())\n        await ws.send(json.dumps({\"type\": \"session.close\"}))  # wait for reply before closing\n\n# asyncio.run(main(your_chunks))"},"warnings":[{"severity":"high","title":"China-only account and KYC","detail":"Sold on Volcengine in mainland China; expect real-name verification (personal ID or business licence) before paid API use, a Chinese-language console and CNY billing. No BytePlus international equivalent was found."},{"severity":"high","title":"Two incompatible protocols","detail":"The 3.0 Seeduplex endpoint uses JSON Realtime-style events; O/SC/2.0 use a custom binary framing on a different path. Code and samples are not interchangeable; migrate event names first, then move asr/tts/dialog config into 'extension'."},{"severity":"medium","title":"Strict real-time pacing","detail":"Sending audio faster or slower than real time triggers server errors, and stopping the uplink without a mute event causes timeouts. File-based tests must sleep 20 ms per 20 ms chunk."},{"severity":"medium","title":"Output audio is the cost driver","detail":"Output audio is 300 CNY per 1M tokens at 25 tokens per second, so talkative agents cost far more than listeners. Token ratios are documented as subject to change; bill on metered tokens, not estimates."},{"severity":"medium","title":"Low default throughput","detail":"60 sessions started per minute and 100k tokens per minute per AppID by default; a busy call centre needs a quota increase arranged with sales in advance."},{"severity":"low","title":"Chinese-first quality","detail":"Docs state only Chinese and English; cloned voices are only reliably good in Chinese."},{"severity":"low","title":"Postpaid billing lag","detail":"Postpaid bills are issued hourly with possible delays of several hours, so keep a balance buffer to avoid suspension."}],"best_for":"Chinese-language consumer voice products, role-play/companion apps and in-car or device assistants targeting mainland China.","open_source":false,"self_hostable":false,"compliance":"Not stated for this API in the docs read.","docs":[{"label":"Realtime 3.0 full-duplex API","url":"https://docs.volcengine.com/docs/DoubaoVoice/endtoend-realtime-voice-full-duplex-version?lang=zh"},{"label":"Access must-read (3.0 protocol)","url":"https://docs.volcengine.com/docs/DoubaoVoice/access-mustread?lang=zh"},{"label":"Legacy end-to-end API","url":"https://docs.volcengine.com/docs/6561/1594356"},{"label":"Billing","url":"https://docs.volcengine.com/docs/6561/1359370?lang=zh"}],"sources":["https://docs.volcengine.com/docs/6561/1594356","https://docs.volcengine.com/docs/DoubaoVoice/access-mustread?lang=zh","https://docs.volcengine.com/docs/DoubaoVoice/endtoend-realtime-voice-full-duplex-version?lang=zh","https://docs.volcengine.com/docs/6561/1359370?lang=zh"],"confidence":"medium","unverified":"Free-quota size and resource-pack prices; whether the worked-example rates apply identically to 3.0 Seeduplex and to 2.0; exact session.create payload fields beyond session.model; KYC rules for foreign users.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Doubao Realtime","facts":{"latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":false,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":true,"context_tokens":null,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":11.2676,"price_audio_out_per_1m":42.2535,"price_per_min":null,"_notes":"Prices converted from 80 / 300 CNY per 1M tokens at 7.1 CNY per USD. Free quota exists but size not documented. Mainland China only with real-name verification. Default 60 session starts per minute and 100k TPM per AppID. Connection released after 10 min idle. Older 2.0 models have 12K context."}},{"id":"zhipu-glm-realtime","name":"GLM-Realtime","vendor":"Zhipu AI (bigmodel.cn)","category":"speech-to-speech","summary":"Zhipu's realtime voice (and passive video) model over an OpenAI-Realtime-style WebSocket, billed simply per minute. Good for Chinese-market voice and video-call assistants; short memory makes it a poor fit for long sessions.","status":"GA","models":[{"name":"glm-realtime-flash","status":"GA","notes":"9B model per docs."},{"name":"glm-realtime-air","status":"GA","notes":"32B model per docs."},{"name":"glm-realtime","status":"GA","notes":"Default value of session.model; no separate price listed."}],"transports":["WebSocket"],"audio":{"input":"wav or pcm (pcm16 = 16 kHz, pcm24 = 24 kHz), mono 16-bit","output":"PCM 24 kHz mono 16-bit"},"languages":"Multilingual with automatic language detection (no list published); replies in the user's language.","voices":"tongtong (default), xiaochen, female-tianmei, female-shaonv, male-qn-daxuesheng, male-qn-jingying, lovely_girl","latency":"Not published.","features":["server VAD or client VAD","interruption (interrupt_response, response.cancel)","function calling (voice calls only)","built-in web search (auto_search)","passive video mode (video_passive)","near/far-field noise reduction","greeting config","singing"],"pricing":{"model":"per-minute","items":[{"what":"glm-realtime-flash audio call","price":"0.18 CNY","unit":"per minute","notes":""},{"what":"glm-realtime-flash video call","price":"1.2 CNY","unit":"per minute","notes":""},{"what":"glm-realtime-air audio call","price":"0.3 CNY","unit":"per minute","notes":""},{"what":"glm-realtime-air video call","price":"2.1 CNY","unit":"per minute","notes":""}],"audio_token_rate":"Not token-billed.","est_per_minute_usd":{"low":0.025,"high":0.3,"basis":"Conversion at about 7.1 CNY per USD: flash audio about $0.025, air audio about $0.042, air video about $0.30 per minute."},"free_tier":"Not stated on the GLM-Realtime page.","source":"https://docs.bigmodel.cn/cn/guide/models/sound-and-video/glm-realtime"},"limits":["Context 8K for audio calls (about 20 turns per docs) and 32K for video calls","Conversation memory up to about 2 minutes","max_response_output_tokens up to 1024","Client VAD mode: max 30 s per upload; send at most 50 messages per second (100 ms frames recommended)","Concurrency by account tier: V0 5, V1 10, V2 15, V3 20"],"regions":"Chinese mainland (open.bigmodel.cn). Not confirmed on the international z.ai platform (its GLM-Realtime doc URL returned 404).","setup":{"steps":["Register on open.bigmodel.cn and create an API key (real-name verification may be required).","Connect to the realtime WebSocket with Authorization: Bearer <key> (a JWT also works).","Send session.update with session.model set to glm-realtime-flash or glm-realtime-air.","Stream input_audio_buffer.append (pcm16) and play response.audio.delta (24 kHz PCM)."],"endpoint":"wss://open.bigmodel.cn/api/paas/v4/realtime","auth":"Authorization header with API key or JWT","snippet_lang":"javascript","snippet":"// npm i ws\nimport WebSocket from \"ws\";\n\nconst ws = new WebSocket(\"wss://open.bigmodel.cn/api/paas/v4/realtime\", {\n  headers: { Authorization: `Bearer ${process.env.ZHIPU_API_KEY}` },\n});\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"session.update\",\n    session: {\n      model: \"glm-realtime-flash\",        // or glm-realtime-air\n      modalities: [\"text\", \"audio\"],\n      instructions: \"You are a helpful voice assistant.\",\n      voice: \"tongtong\",\n      input_audio_format: \"pcm16\",        // 16 kHz mono 16-bit\n      output_audio_format: \"pcm\",         // 24 kHz mono 16-bit\n      beta_fields: { chat_mode: \"audio\" },\n    },\n  }));\n});\n\n// ~100 ms frames, at most 50 messages per second\nexport function sendPcm(buf) {\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: buf.toString(\"base64\") }));\n}\n\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"response.audio.delta\") play(Buffer.from(ev.delta, \"base64\"));\n  if (ev.type === \"error\") console.error(ev);\n});\nfunction play(pcm) {}"},"warnings":[{"severity":"high","title":"Very short memory","detail":"Audio sessions get an 8K context (about 20 turns) and roughly 2 minutes of conversation memory; long calls will forget earlier details. Re-inject key facts via instructions."},{"severity":"medium","title":"China platform only","detail":"Documented on open.bigmodel.cn with CNY pricing; international z.ai availability is not confirmed. Expect Chinese account requirements and mainland data processing."},{"severity":"medium","title":"Low concurrency on new accounts","detail":"Starting tier allows 5 concurrent sessions; higher tiers depend on account level."},{"severity":"medium","title":"Video is 6-7x the audio price","detail":"Video mode is 1.2 to 2.1 CNY per minute and needs at least one image uploaded before creating a response, or it errors."},{"severity":"low","title":"Short replies","detail":"max_response_output_tokens is capped at 1024, so long answers get cut."},{"severity":"low","title":"Tools only in voice mode","detail":"Function calling is documented as voice-call only, not video calls."}],"best_for":"Chinese-market voice or video-call assistants with short interactions and predictable per-minute cost.","open_source":false,"self_hostable":false,"compliance":"Not stated on the page read.","docs":[{"label":"GLM-Realtime guide","url":"https://docs.bigmodel.cn/cn/guide/models/sound-and-video/glm-realtime"}],"sources":["https://docs.bigmodel.cn/cn/guide/models/sound-and-video/glm-realtime"],"confidence":"medium","unverified":"International (z.ai) availability, free tier, real-name rules for foreigners, latency.","cat":"voice","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":5,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":false,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":true,"voices":7,"voice_cloning":null,"context_tokens":8000,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.0254,"_notes":"Price is glm-realtime-flash audio, 0.18 CNY/min converted at 7.1 CNY per USD (air 0.3 CNY/min; video 1.2 to 2.1 CNY/min). Concurrency 5 at account tier V0, up to 20 at V3. Context 8K for audio (about 2 min memory), 32K for video. China platform only."}},{"id":"stepfun-realtime","name":"StepAudio Realtime","vendor":"StepFun","category":"speech-to-speech","summary":"StepFun's end-to-end voice models (StepAudio 2.5 Realtime, Step-Audio 2) on an OpenAI-Realtime-style WebSocket, with voice cloning and paralinguistic cues like laughs and sighs. Suits Chinese-first companion and role-play apps.","status":"GA","models":[{"name":"stepaudio-2.5-realtime","status":"GA","notes":"Released May 2026 per press; roleplay-focused RLHF. A news article used the string step-2.5-realtime, docs use stepaudio-2.5-realtime."},{"name":"step-audio-2","status":"GA","notes":"Previous generation."},{"name":"step-audio-2-mini","status":"GA","notes":"Listed in the realtime guide as a model value; not on the price table. Open weights exist (see open-model entry)."},{"name":"step-1o-audio","status":"GA","notes":"Older model, still priced."}],"transports":["WebSocket"],"audio":{"input":"pcm16 (sample rate not stated on the model page)","output":"pcm16"},"languages":"Not listed; docs and examples are Chinese.","voices":"Preset voices such as linjiajiejie; voice cloning via uploaded reference audio returns a custom voice id.","latency":"Not published.","features":["server VAD","streaming audio deltas","voice cloning","persona instructions","paralinguistic output (laughs, sighs)"],"pricing":{"model":"per-token","items":[{"what":"stepaudio-2.5-realtime","price":"$1.50 in (cache miss) / $0.30 in (cache hit) / $10.00 out","unit":"per 1M tokens","notes":"China site: 10 / 2 / 70 CNY."},{"what":"step-audio-2","price":"$1.43 / $0.29 / $10.00","unit":"per 1M tokens","notes":""},{"what":"step-1o-audio","price":"$3.57 / $0.71 / $8.57","unit":"per 1M tokens","notes":""}],"audio_token_rate":"Not published, so per-minute cost cannot be derived from docs.","est_per_minute_usd":{"low":null,"high":null,"basis":"Audio tokens per second are not documented; measure usage on a test call."},"free_tier":"Not confirmed.","source":"https://platform.stepfun.ai/docs/en/guides/pricing/details"},"limits":["Account rate tiers on the open platform range from V0 (under $15 top-up: 5 concurrency, 100 RPM) to V4 ($1,500+: 130 concurrency, 2,600 RPM); not stated whether these apply to realtime","Session and context limits not documented on the model page"],"regions":"China platform (platform.stepfun.com, CNY) and a USD-priced platform at platform.stepfun.ai; realtime availability on the .ai platform not confirmed.","setup":{"steps":["Create an account on the StepFun open platform and an API key.","Connect to /v1/realtime with ?model=stepaudio-2.5-realtime and Authorization: Bearer <key>.","Send session.update (instructions, voice, pcm16 formats, server_vad), stream audio, play response.audio.delta."],"endpoint":"wss://api.stepfun.com/v1/realtime?model=stepaudio-2.5-realtime","auth":"Authorization: Bearer <STEPFUN_API_KEY>","snippet_lang":"javascript","snippet":"// npm i ws   (China platform host shown; check the console for the international host)\nimport WebSocket from \"ws\";\n\nconst ws = new WebSocket(\"wss://api.stepfun.com/v1/realtime?model=stepaudio-2.5-realtime\", {\n  headers: { Authorization: `Bearer ${process.env.STEPFUN_API_KEY}` },\n});\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"session.update\",\n    session: {\n      modalities: [\"text\", \"audio\"],\n      instructions: \"You are a warm, concise assistant.\",\n      voice: \"linjiajiejie\",\n      input_audio_format: \"pcm16\",\n      output_audio_format: \"pcm16\",\n      turn_detection: { type: \"server_vad\", prefix_padding_ms: 500 },\n    },\n  }));\n});\n\nexport function sendPcm(buf) {\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: buf.toString(\"base64\") }));\n}\n\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"response.audio.delta\") play(Buffer.from(ev.delta, \"base64\"));\n});\nfunction play(pcm) {}"},"warnings":[{"severity":"high","title":"Cost per minute is unknowable from docs","detail":"Billing is per token but the audio tokens-per-second rate is not published. Run a metered test call before quoting customers."},{"severity":"medium","title":"Model string inconsistency","detail":"Docs use stepaudio-2.5-realtime, press used step-2.5-realtime, and third-party lists vary. Check the console model list before hardcoding."},{"severity":"medium","title":"Chinese-first docs","detail":"The detailed realtime docs are Chinese-only; language support, function calling and session limits are not documented."},{"severity":"medium","title":"International availability unclear","detail":"A USD price list exists on platform.stepfun.ai but it does not say realtime models are available there; the documented endpoint is api.stepfun.com."},{"severity":"low","title":"Concurrency tied to top-up","detail":"Platform concurrency scales with cumulative spend (5 concurrent at the lowest tier), if those tiers apply to realtime."}],"best_for":"Expressive Chinese-language companion or character voice apps.","open_source":false,"self_hostable":false,"compliance":"Not stated.","docs":[{"label":"StepAudio 2.5 Realtime (zh)","url":"https://platform.stepfun.com/docs/zh/guides/models/stepaudio-2.5-realtime"},{"label":"Realtime developer guide (zh)","url":"https://platform.stepfun.com/docs/zh/guides/developer/realtime"},{"label":"Pricing (en)","url":"https://platform.stepfun.ai/docs/en/guides/pricing/details"}],"sources":["https://platform.stepfun.com/docs/zh/guides/models/stepaudio-2.5-realtime","https://platform.stepfun.ai/docs/en/guides/pricing/details","https://www.marktechpost.com/2026/05/24/stepfun-releases-stepaudio-2-5-realtime-an-end-to-end-voice-model-with-roleplay-specific-rlhf-and-paralinguistic-comprehension/"],"confidence":"medium","unverified":"Audio token rate, sample rates, languages, function calling, international endpoint, free tier.","cat":"voice","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":null,"image_input":null,"voices":null,"voice_cloning":true,"context_tokens":null,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":1.5,"price_audio_out_per_1m":10,"price_per_min":null,"_notes":"Prices are stepaudio-2.5-realtime per 1M tokens (cache-hit input $0.30); audio tokens per second not published. Platform concurrency tiers (5 at V0) may not apply to realtime. Docs are Chinese-first."}},{"id":"hume-evi","name":"Hume EVI (Empathic Voice Interface)","vendor":"Hume AI","category":"speech-to-speech","summary":"Emotion-aware voice interface that measures vocal expression and replies with expressive speech, optionally using a supplemental LLM. Hume is shutting EVI and its TTS API down on November 13, 2026, so it is only relevant for migrating existing users.","status":"Deprecated","models":[{"name":"EVI 3","status":"GA","notes":"English only; quick responses; supplemental LLM optional. Ends 2026-11-13."},{"name":"EVI 4-mini","status":"GA","notes":"11 languages; supplemental LLM required; no quick responses. Released Oct 2025. Ends 2026-11-13."}],"transports":["WebSocket"],"audio":{"input":"WebM (browser) or linear16 PCM (e.g. 44.1 kHz mono, declared in session_settings); no mu-law","output":"base64 WAV in audio_output messages"},"languages":"EVI 3: English. EVI 4-mini: English, Japanese, Korean, Spanish, French, Portuguese, Italian, German, Russian, Hindi, Arabic.","voices":"Hume voice library and custom voices via configs.","latency":"No figure on the overview page.","features":["expression (prosody) measures per sentence","tone-aware end-of-turn detection","interruption handling","supplemental LLMs (Anthropic, OpenAI, Google, Fireworks and others)","configs and chat history resume"],"pricing":{"model":"subscription","items":[{"what":"Free","price":"$0","unit":"per month","notes":"5 EVI minutes (third-party data)"},{"what":"Starter","price":"$3","unit":"per month","notes":"40 EVI minutes, overage $0.07/min (third-party data)"},{"what":"Creator","price":"$14","unit":"per month","notes":"200 minutes, overage $0.07/min (third-party data)"},{"what":"Pro","price":"$70","unit":"per month","notes":"1,200 minutes, overage $0.06/min (third-party data)"},{"what":"Scale","price":"$200","unit":"per month","notes":"5,000 minutes, overage $0.05/min (third-party data)"},{"what":"Business","price":"$500","unit":"per month","notes":"12,500 minutes, overage $0.04/min (third-party data)"}],"audio_token_rate":"Not token-billed.","est_per_minute_usd":{"low":0.04,"high":0.07,"basis":"Third-party reported EVI 3 overage rates; EVI 4-mini reportedly about half. Supplemental LLM cost may be extra."},"free_tier":"5 EVI minutes per month (third-party data).","source":"https://fish.audio/vs/pricing/hume-ai/ (third party; hume.ai/pricing did not render EVI prices)"},"limits":["Access ends November 13, 2026 at 12:01 a.m. EST; account data permanently deleted after that","Max session duration 30 minutes","Max WebSocket message 16 MB","HTTP rate limit 100 requests per second","Concurrent sessions set by subscription tier"],"regions":"Not stated.","setup":{"steps":["Do not start new projects: the API shuts down on 2026-11-13.","Existing users: export any chat data you need before the deadline and plan a migration (Hume points to Real World VoiceEQ for comparing alternatives).","Connect to the EVI chat WebSocket with api_key (server) or access_token (browser) and a config_id."],"endpoint":"wss://api.hume.ai/v0/evi/chat","auth":"api_key or access_token query parameter (REST uses X-HUME-API-KEY header)","snippet_lang":"javascript","snippet":"// EVI is being shut down on 2026-11-13. Only for existing integrations.\n// npm i ws. Use an access token (not the API key) from browsers.\nimport WebSocket from \"ws\";\n\nconst url = `wss://api.hume.ai/v0/evi/chat?api_key=${process.env.HUME_API_KEY}&config_id=${process.env.HUME_CONFIG_ID}`;\nconst ws = new WebSocket(url);\n\nws.on(\"open\", () => {\n  // Raw PCM must be declared; WebM from browsers needs no settings.\n  ws.send(JSON.stringify({\n    type: \"session_settings\",\n    audio: { encoding: \"linear16\", sample_rate: 44100, channels: 1 },\n  }));\n});\n\nexport function sendPcm(buf) {\n  ws.send(JSON.stringify({ type: \"audio_input\", data: buf.toString(\"base64\") }));\n}\n\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"audio_output\") playWav(Buffer.from(ev.data, \"base64\")); // base64 WAV\n  if (ev.type === \"user_message\") console.log(ev.message?.content, ev.models?.prosody);\n});\nfunction playWav(wav) {}"},"warnings":[{"severity":"high","title":"Shutting down November 13, 2026","detail":"Hume's docs and changelog (notice dated 2026-10-02) say the TTS and EVI APIs end at 12:01 a.m. EST on November 13, 2026 and all account data is then permanently deleted. Do not build on it."},{"severity":"high","title":"Export data now","detail":"Chat histories, configs and custom voices are deleted after the cutoff; export transcripts and anything needed for compliance before then."},{"severity":"medium","title":"Pricing not verifiable from Hume","detail":"Hume's pricing page no longer shows EVI rates; figures here come from third-party comparison pages and may be stale."},{"severity":"medium","title":"EVI 4-mini needs a supplemental LLM","detail":"4-mini requires a supplemental LLM (third-party sources say LLM cost is billed separately), so the per-minute rate is not the full cost."},{"severity":"medium","title":"30-minute session cap","detail":"Sessions end at 30 minutes; long calls need resume via chat group ids."},{"severity":"low","title":"EVI 3 is English only","detail":"Only EVI 4-mini is multilingual."}],"best_for":"Nothing new. Existing integrations should migrate before 2026-11-13.","open_source":false,"self_hostable":false,"compliance":"Not re-verified; irrelevant after shutdown.","docs":[{"label":"EVI overview (sunset notice)","url":"https://dev.hume.ai/docs/speech-to-speech-evi/overview"},{"label":"Changelog","url":"https://dev.hume.ai/changelog"},{"label":"EVI chat WebSocket reference","url":"https://dev.hume.ai/reference/speech-to-speech-evi/chat"}],"sources":["https://dev.hume.ai/docs/speech-to-speech-evi/overview","https://dev.hume.ai/docs/introduction","https://dev.hume.ai/changelog","https://dev.hume.ai/docs/speech-to-speech-evi/guides/audio","https://fish.audio/vs/pricing/hume-ai/","https://usagepricing.com/blueprint/hume-ai"],"confidence":"medium","unverified":"All prices (third-party only); exact session_settings field names in the snippet; whether supplemental LLM usage is billed separately.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Hume EVI","facts":{"latency_ms":null,"languages":11,"max_session_min":30,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":null,"image_input":null,"voices":null,"voice_cloning":true,"context_tokens":null,"byo_llm":true,"native_s2s":true,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.07,"_notes":"Shutting down November 13, 2026. 11 languages on EVI 4-mini (EVI 3 English only). Price is third-party reported Starter overage ($0.04 to $0.07 by plan). Free plan 5 min/month (third-party). Supplemental LLM optional on EVI 3, required on EVI 4-mini."}},{"id":"elevenlabs-agents","name":"ElevenAgents (formerly Conversational AI / Agents Platform)","vendor":"ElevenLabs","category":"voice-agent","summary":"A hosted voice-agent platform: ElevenLabs speech recognition, a proprietary turn-taking model, your choice of LLM and ElevenLabs voices, with telephony, RAG and tools. Best for teams that want premium voices and a full no-code/low-code agent builder.","status":"GA","models":[{"name":"Platform (STT + turn-taking + LLM + ElevenLabs TTS)","status":"GA","notes":"Cascaded pipeline, not a single speech-to-speech model."},{"name":"LLM choices","status":"GA","notes":"OpenAI (GPT-4o mini to GPT-6.1 family), Anthropic (Haiku 4.5 to Opus 5.5), Google Gemini 2.5 to 3.8 Flash, plus ElevenLabs-hosted open models (GLM 5.2, Qwen3.6-35B-A3B, Qwen3.5-397B-A17B, DeepSeek Flash 4.1) or a custom LLM endpoint."}],"transports":["WebSocket","SIP","Twilio","Web/mobile SDKs","Embeddable widget"],"audio":{"input":"pcm_8000/16000/22050/24000/44100/48000 or ulaw_8000 (negotiated per agent)","output":"Same enum as input"},"languages":"Docs give both '31 languages' and '70+ languages' for the voice layer (inconsistent).","voices":"5k+ voices in the ElevenLabs library, plus voice clones.","latency":"Described as low-latency; no figure on the overview page.","features":["tool calling (server and client tools)","MCP tools with approval","RAG knowledge base","proprietary turn-taking model","interruptions and timeouts config","workflow builder","multimodal text+voice","burst concurrency","data residency endpoints"],"pricing":{"model":"per-minute","items":[{"what":"Agent call minutes (all plans)","price":"$0.08","unit":"per minute","notes":"Down from $0.10 on 2026-05-07. Billed on connection time; silence over 10 s billed at 5% of the rate."},{"what":"Burst calls above concurrency limit","price":"$0.16","unit":"per minute","notes":"Up to 3x the plan concurrency."},{"what":"Text messages","price":"$0.003","unit":"per message","notes":""},{"what":"LLM","price":"provider list price","unit":"per 1M tokens","notes":"Passed through with no markup; select Gemini and Claude models billed at Vertex regional rates (+10% for US/EU)."},{"what":"Plans (included minutes / concurrency)","price":"Free 15/4, Starter $6 75/6, Creator $22 275/10, Pro $99 1,238/20, Scale $299 3,738/30, Business $990 12,375/40","unit":"per month","notes":"Enterprise custom."},{"what":"Telephony","price":"at cost","unit":"","notes":"Twilio or SIP carrier bills you directly."}],"audio_token_rate":"Not token-billed (LLM portion is token-billed).","est_per_minute_usd":{"low":0.08,"high":0.13,"basis":"$0.08 platform rate plus LLM pass-through; a third-party guide puts LLM at about $0.0005/min (GPT-5 Nano) to $0.045/min (GPT-5.5). Telephony and burst pricing extra."},"free_tier":"15 call minutes per month, 4 concurrent calls.","source":"https://elevenlabs.io/pricing/agents"},"limits":["Concurrency 4 (Free) to 40 (Business) concurrent calls; burst to 3x at double price","Some LLMs are unavailable when EU data residency is enabled","Call time measured on connection duration, not speech time"],"regions":"Default global, plus wss://api.us.elevenlabs.io, EU (api.eu.residency.elevenlabs.io), India (api.in.residency.elevenlabs.io) and Singapore (api.sg.residency.elevenlabs.io) residency hosts.","setup":{"steps":["Create an ElevenLabs account and build an agent in the dashboard (prompt, LLM, voice, tools, audio formats).","For public agents connect with ?agent_id; for private agents generate a signed URL server-side.","Stream base64 user_audio_chunk messages, play 'audio' events, answer 'ping' with 'pong'.","For phone calls attach a Twilio number or SIP trunk in the dashboard."],"endpoint":"wss://api.elevenlabs.io/v1/convai/conversation?agent_id=<agent_id>","auth":"agent_id query parameter for public agents; signed URL / token for private agents (xi-api-key for REST)","snippet_lang":"javascript","snippet":"// npm i ws. Create the agent in the dashboard first; set its audio formats to pcm_16000.\n// Public agents: agent_id in the URL. Private agents: mint a signed URL on your server.\nimport WebSocket from \"ws\";\n\nconst ws = new WebSocket(\n  `wss://api.elevenlabs.io/v1/convai/conversation?agent_id=${process.env.ELEVEN_AGENT_ID}`\n  // EU residency: wss://api.eu.residency.elevenlabs.io/...\n);\n\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  switch (ev.type) {\n    case \"conversation_initiation_metadata\":\n      console.log(ev.conversation_initiation_metadata_event); // negotiated audio formats\n      break;\n    case \"audio\":\n      play(Buffer.from(ev.audio_event.audio_base_64, \"base64\"));\n      break;\n    case \"ping\":\n      ws.send(JSON.stringify({ type: \"pong\", event_id: ev.ping_event.event_id }));\n      break;\n  }\n});\n\n// 16 kHz mono 16-bit PCM from your mic\nexport function sendPcm(buf) {\n  ws.send(JSON.stringify({ user_audio_chunk: buf.toString(\"base64\") }));\n}\nfunction play(pcm) {}"},"warnings":[{"severity":"high","title":"LLM is not in the $0.08","detail":"LLM tokens are billed on top at provider rates (Gemini/Claude via Vertex regional pricing are 10% higher for US/EU). Large models can add roughly half the platform rate again."},{"severity":"high","title":"Connection time, not talk time","detail":"Billing counts the whole connection; a tab left open or a call never hung up keeps billing (silence over 10 s at 5%). Set max conversation duration and inactivity timeouts."},{"severity":"medium","title":"Burst pricing doubles cost","detail":"Calls above plan concurrency are allowed up to 3x but charged $0.16/min. Spiky outbound campaigns can silently double the bill."},{"severity":"medium","title":"Residency limits model choice","detail":"With EU data residency some LLMs are unavailable; check model availability per region before committing."},{"severity":"medium","title":"Old rate cards everywhere","detail":"Prices fell in May 2026 ($0.10 to $0.08) and included minutes differ across third-party guides; use elevenlabs.io/pricing/agents."},{"severity":"low","title":"Telephony billed separately","detail":"Twilio/SIP minutes are paid to the carrier; ElevenLabs adds no telephony fee."},{"severity":"low","title":"HIPAA needs Enterprise","detail":"BAAs are listed under Enterprise only."}],"best_for":"Customer-facing agents where voice quality and a managed builder (RAG, tools, telephony, evals) matter more than raw cost.","open_source":false,"self_hostable":false,"compliance":"Enterprise: DPA/SLA, BAA for HIPAA, custom SSO; regional data residency hosts (EU, India, Singapore). SOC 2/GDPR not re-verified this session.","docs":[{"label":"ElevenAgents overview","url":"https://elevenlabs.io/docs/agents-platform/overview"},{"label":"WebSocket API","url":"https://elevenlabs.io/docs/agents-platform/api-reference/agents-platform/websocket"},{"label":"LLM options and pass-through","url":"https://elevenlabs.io/docs/agents-platform/customization/llm"},{"label":"Pricing","url":"https://elevenlabs.io/pricing/agents"}],"sources":["https://elevenlabs.io/pricing/agents","https://elevenlabs.io/docs/help-center/product/eleven-agents/how-much-does-eleven-agents-cost","https://elevenlabs.io/blog/weve-lowered-api-agents-pricing-and-introduced-pay-as-you-go","https://elevenlabs.io/docs/agents-platform/customization/llm","https://elevenlabs.io/docs/agents-platform/api-reference/agents-platform/websocket","https://www.getmacha.com/blog/elevenlabs-agents-pricing-explained"],"confidence":"high","unverified":"Per-minute LLM cost examples come from a third-party guide; WebRTC transport availability not confirmed on the pages read; supported language count is inconsistent in the docs.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"ElevenLabs Agents","facts":{"latency_ms":null,"languages":31,"max_session_min":null,"concurrency":6,"free_tier":true,"free_credit_usd":null,"hipaa":true,"soc2":null,"gdpr_eu":true,"webrtc":null,"websocket":true,"sip":true,"open_weights":false,"function_calling":true,"image_input":null,"voices":5000,"voice_cloning":true,"context_tokens":null,"byo_llm":true,"native_s2s":false,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.08,"_notes":"Concurrency is the Starter plan (Free 4, Business 40, burst to 3x at $0.16/min). $0.08/min excludes LLM tokens. Free plan 15 min/month. BAA on Enterprise only. Docs also mention 70+ languages for the voice layer.","_added":["languages: https://elevenlabs.io/docs/agents-platform/customization/language"]}},{"id":"deepgram-voice-agent","name":"Deepgram Voice Agent API","vendor":"Deepgram","category":"voice-agent","summary":"One WebSocket that runs Deepgram speech recognition, a managed or bring-your-own LLM, and Deepgram Aura or third-party TTS, billed per connected minute. Good for developers who want a simple flat rate with the LLM included on Standard models.","status":"GA","models":[{"name":"STT: Deepgram Flux / Nova (e.g. flux-general-en, flux-general-multi)","status":"GA","notes":""},{"name":"LLM Standard tier","status":"GA","notes":"e.g. gpt-5-mini, gpt-5.4-mini, gpt-4o-mini, claude-haiku-4-5, gemini-3.5-flash, gemini-2.5-flash, nemotron-3-nano, groq openai/gpt-oss-20b (BYO)."},{"name":"LLM Advanced tier","status":"GA","notes":"e.g. gpt-5.5, gpt-5.4, gpt-5, gpt-4.1, gpt-4o, claude-sonnet-5, claude-sonnet-4-6, gemini-3-pro-preview."},{"name":"TTS: Deepgram Aura-2 / Flux TTS, or eleven_labs, cartesia, open_ai, aws_polly","status":"GA","notes":"Non-Deepgram TTS requires your own endpoint/key (BYO TTS tiers)."}],"transports":["WebSocket"],"audio":{"input":"linear16 default at 16 kHz (other encodings/rates configurable)","output":"linear16 (e.g. 24 kHz), optional container; other encodings configurable"},"languages":"English by default; multilingual via Nova 'multi' or flux-general-multi with language hints, and multilingual TTS providers.","voices":"Deepgram Aura-2 voices (e.g. aura-2-thalia-en) or third-party voices.","latency":"Not stated on the pages read; the API reports per-turn latency metrics (time to first token, TTS time to first byte).","features":["function calling (client or server side)","barge-in (UserStartedSpeaking)","greeting","LLM fallback","reusable stored agent configurations","browser agent SDK and widget","mid-session updates of prompt/think/speak"],"pricing":{"model":"per-minute","items":[{"what":"Standard","price":"$0.075 PAYG / $0.068 Growth","unit":"per minute","notes":"Includes Deepgram STT, a Standard-tier managed LLM and Deepgram TTS."},{"what":"Standard - BYO TTS","price":"$0.065 / $0.051","unit":"per minute","notes":""},{"what":"Custom - BYO LLM","price":"$0.065 / $0.059","unit":"per minute","notes":"You also pay your LLM provider."},{"what":"Custom - BYO LLM + TTS","price":"$0.050 / $0.041","unit":"per minute","notes":"You also pay LLM and TTS providers."},{"what":"Advanced","price":"$0.163 / $0.146","unit":"per minute","notes":"Advanced-tier LLMs."},{"what":"Advanced - BYO TTS","price":"$0.122 / $0.110","unit":"per minute","notes":""}],"audio_token_rate":"Not token-billed; billed on WebSocket connection time.","est_per_minute_usd":{"low":0.041,"high":0.163,"basis":"Official tier rates; BYO tiers add your own LLM/TTS bills on top."},"free_tier":"$200 one-time credit for new accounts.","source":"https://deepgram.com/pricing"},"limits":["Concurrency: up to 45 WebSocket agent sessions on Pay As You Go, up to 60 on Growth","Billing runs for the whole WebSocket connection"],"regions":"Hosted API (agent.deepgram.com); Deepgram also sells self-hosted deployments (not verified for the agent API).","setup":{"steps":["Create a Deepgram account (includes $200 credit) and an API key.","Open wss://agent.deepgram.com/v1/agent/converse with Authorization: Token <key>.","Send a Settings message (audio formats, listen/think/speak providers, prompt, greeting).","Stream raw audio as binary frames, play binary audio back, send KeepAlive when idle, stop playback on UserStartedSpeaking."],"endpoint":"wss://agent.deepgram.com/v1/agent/converse","auth":"Authorization header (Token <API key>, or a short-lived token for browsers)","snippet_lang":"javascript","snippet":"// npm i ws\nimport WebSocket from \"ws\";\n\nconst ws = new WebSocket(\"wss://agent.deepgram.com/v1/agent/converse\", {\n  headers: { Authorization: `Token ${process.env.DEEPGRAM_API_KEY}` },\n});\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"Settings\",\n    audio: {\n      input: { encoding: \"linear16\", sample_rate: 16000 },\n      output: { encoding: \"linear16\", sample_rate: 24000, container: \"none\" },\n    },\n    agent: {\n      listen: { provider: { type: \"deepgram\", model: \"flux-general-en\", version: \"v2\" } },\n      think: { provider: { type: \"open_ai\", model: \"gpt-4o-mini\" }, // a \"Standard\" tier model\n               prompt: \"You are a helpful assistant. Keep answers short.\" },\n      speak: { provider: { type: \"deepgram\", model: \"aura-2-thalia-en\" } },\n      greeting: \"Hi! How can I help?\",\n    },\n  }));\n  setInterval(() => ws.send(JSON.stringify({ type: \"KeepAlive\" })), 5000);\n});\n\n// Send raw linear16 mic audio as binary frames\nexport function sendPcm(buf) { ws.send(buf); }\n\nws.on(\"message\", (data, isBinary) => {\n  if (isBinary) return play(data);          // agent speech (raw PCM)\n  const ev = JSON.parse(data.toString());\n  if (ev.type === \"UserStartedSpeaking\") stopPlayback(); // barge-in\n  if (ev.type === \"Error\") console.error(ev);\n});\nfunction play(pcm) {} function stopPlayback() {}"},"warnings":[{"severity":"high","title":"Model choice changes the tier","detail":"Picking an Advanced-tier LLM (e.g. GPT-5, Claude Sonnet) moves the session to $0.163/min, more than double Standard. The tier follows the model, so a config change can double cost."},{"severity":"high","title":"Connection time is billed","detail":"The meter runs while the WebSocket is open, including silence and hold. Close sockets promptly and avoid idle KeepAlive loops."},{"severity":"medium","title":"BYO tiers are not all-in","detail":"BYO LLM/TTS rates exclude what you pay OpenAI, Bedrock, Groq, ElevenLabs etc. Bedrock and Groq always require your own endpoint and credentials."},{"severity":"medium","title":"Promotional rates expired","detail":"Third-party pages still quote promo or older rates (e.g. $0.08 Standard, $0.056 BYO LLM); some promos ended 2026-09-12. Use deepgram.com/pricing."},{"severity":"medium","title":"Concurrency ceiling","detail":"45 concurrent sessions on PAYG and 60 on Growth; larger contact centres need Enterprise."},{"severity":"low","title":"Deprecated fields","detail":"agent.language is deprecated; set language on listen.provider and speak.provider. Some listed models are already marked deprecated."},{"severity":"low","title":"Cascaded, not native speech-to-speech","detail":"STT, LLM and TTS are separate stages, so paralinguistic cues (tone, laughter) are not passed to the LLM."}],"best_for":"Developers who want one WebSocket, predictable per-minute cost and strong telephony-grade STT, with an option to swap in their own LLM or TTS.","open_source":false,"self_hostable":false,"compliance":"Not re-verified this session (Deepgram markets SOC 2 and HIPAA readiness; check the trust page).","docs":[{"label":"Voice Agent getting started","url":"https://developers.deepgram.com/docs/voice-agent"},{"label":"Configure the agent","url":"https://developers.deepgram.com/docs/configure-voice-agent"},{"label":"LLM models and tiers","url":"https://developers.deepgram.com/docs/voice-agent-llm-models"},{"label":"API reference","url":"https://developers.deepgram.com/reference/voice-agent/voice-agent"}],"sources":["https://deepgram.com/pricing","https://developers.deepgram.com/docs/voice-agent-llm-models.md","https://developers.deepgram.com/docs/configure-voice-agent.md","https://developers.deepgram.com/reference/voice-agent/voice-agent.md"],"confidence":"high","unverified":"Latency figures; compliance certifications; Growth plan entry requirements; whether 'Token' vs 'Bearer' is required for every key type.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Deepgram Voice Agent","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":45,"free_tier":true,"free_credit_usd":200,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":null,"context_tokens":null,"byo_llm":true,"native_s2s":false,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.075,"_notes":"Price is Standard tier PAYG (BYO LLM+TTS $0.050, Advanced $0.163). Concurrency 45 on PAYG, 60 on Growth. English by default, multilingual via multi models. Deepgram markets SOC 2 and HIPAA readiness, not verified."}},{"id":"ultravox","name":"Ultravox Realtime","vendor":"Ultravox (formerly Fixie.ai)","category":"voice-agent","summary":"Hosted voice agents built on Ultravox's audio-native model, which listens to speech directly (no separate transcription) and answers through TTS voices, at a flat $0.05 per minute. A good default for cost-sensitive phone and web agents; the speech-understanding model weights are open.","status":"GA","models":[{"name":"ultravox-v0.7","status":"GA","notes":"Default since Dec 2025; built on GLM 4.6. Open weights (MIT) for the audio adapter."},{"name":"ultravox-v0.6 / fixie-ai/ultravox-llama3.3-70b","status":"GA","notes":"Legacy (Llama 3.3 70B); must be set explicitly."},{"name":"Qwen3-based Ultravox","status":"Deprecated","notes":"Phased out December 2025."}],"transports":["WebRTC","WebSocket","SIP","Twilio","Telnyx","Plivo","Exotel"],"audio":{"input":"Raw PCM s16le for server WebSocket (sample rate set in the call medium)","output":"PCM for server WebSocket; WebRTC handles codecs automatically"},"languages":"Multilingual; no list on the pages read.","voices":"Ultravox voice library plus voice clones (1 on PAYG, 5 custom voices on Pro).","latency":"Vendor says audio-native processing is faster and robust to transcription errors; no figure on the pages read.","features":["audio-native understanding (no ASR stage)","tool calling","RAG corpora (2 PAYG, 20 Pro)","outbound call scheduler (Pro)","call recordings and transcripts","playground (free)"],"pricing":{"model":"per-minute","items":[{"what":"Calls (Pay As You Go and Pro)","price":"$0.05","unit":"per minute","notes":"Rounded to 6-second 'deciminutes' ($0.005 each); no surge pricing."},{"what":"SIP","price":"$0.005 PAYG / $0.0048 Pro","unit":"per minute","notes":""},{"what":"Text/thread usage","price":"$2.00 in (uncached) / $15.00 out","unit":"per 1M tokens","notes":"Cached input not charged; table layout ambiguous about which plans."},{"what":"Pro plan","price":"$100","unit":"per month","notes":"Platform fee; minutes billed on top; no hard concurrency cap."}],"audio_token_rate":"Not token-billed for calls.","est_per_minute_usd":{"low":0.05,"high":0.055,"basis":"$0.05 call rate plus $0.005 SIP when using SIP; LLM and TTS included."},"free_tier":"30 free call minutes; playground calls free.","source":"https://www.ultravox.ai/pricing"},"limits":["Pay As You Go: hard cap of 5 concurrent calls","maxDuration defaults to 1 hour per call","Invoiced monthly in arrears; early invoice at $10 (PAYG) or $100 (Pro) usage thresholds"],"regions":"Not stated on the pages read.","setup":{"steps":["Create an account and API key in the Ultravox console.","POST /api/calls with systemPrompt, model, voice, medium and maxDuration.","Join the returned joinUrl from a browser (WebRTC client SDK) or bridge it to telephony (Twilio, SIP, server WebSocket)."],"endpoint":"POST https://api.ultravox.ai/api/calls (returns joinUrl)","auth":"X-API-Key header","snippet_lang":"javascript","snippet":"// Server side: create a call (Node 18+, global fetch)\nconst res = await fetch(\"https://api.ultravox.ai/api/calls\", {\n  method: \"POST\",\n  headers: { \"X-API-Key\": process.env.ULTRAVOX_API_KEY, \"Content-Type\": \"application/json\" },\n  body: JSON.stringify({\n    systemPrompt: \"You are a friendly receptionist for Acme Dental.\",\n    model: \"ultravox-v0.7\",\n    maxDuration: \"600s\",          // default is 1 hour; cap it to control spend\n    // medium: { serverWebSocket: { inputSampleRate: 16000, outputSampleRate: 16000 } },\n  }),\n});\nconst { joinUrl } = await res.json();\n\n// Browser side (WebRTC is the default medium): npm i ultravox-client\n// import { UltravoxSession } from \"ultravox-client\";\n// const session = new UltravoxSession();\n// session.joinCall(joinUrl);\n// ... later: session.leaveCall();"},"warnings":[{"severity":"high","title":"5-call concurrency on PAYG","detail":"Pay As You Go has a hard cap of 5 concurrent calls; anything beyond a pilot needs Pro ($100/month) or Enterprise."},{"severity":"medium","title":"Model migrations are forced","detail":"Qwen3-based models were retired with about two weeks' notice in December 2025 and the default moved to v0.7. Pin the model string and watch announcements."},{"severity":"medium","title":"Default call length is 1 hour","detail":"maxDuration defaults to 3600 s; a stuck call can run an hour at $0.05/min. Set a shorter maxDuration."},{"severity":"medium","title":"Open weights are not the whole product","detail":"The MIT-licensed v0.7 checkpoint is an audio adapter on GLM-4.6 (a very large MoE model) and outputs text, not speech; self-hosting needs GLM-4.6 serving plus your own TTS."},{"severity":"low","title":"Deletion is permanent","detail":"Call data is retained until you delete it, and deletion cannot be undone; set a retention process for compliance."},{"severity":"low","title":"Token charges for text threads","detail":"Text/thread usage is billed per token ($2 in / $15 out per 1M), separate from call minutes."}],"best_for":"Budget phone agents and web voice assistants that need decent reasoning at a flat, low per-minute price.","open_source":true,"self_hostable":false,"compliance":"Not stated on the pricing or FAQ pages (a trust portal is linked).","docs":[{"label":"Docs","url":"https://docs.ultravox.ai/"},{"label":"Create call API","url":"https://docs.ultravox.ai/api-reference/calls/calls-post"},{"label":"Pricing","url":"https://www.ultravox.ai/pricing"},{"label":"v0.7 announcement","url":"https://www.ultravox.ai/blog/introducing-ultravox-v0-7-the-world-s-smartest-speech-understanding-model"}],"sources":["https://www.ultravox.ai/pricing","https://docs.ultravox.ai/gettingstarted/faq","https://docs.ultravox.ai/api-reference/calls/calls-post","https://www.ultravox.ai/blog/introducing-ultravox-v0-7-the-world-s-smartest-speech-understanding-model","https://huggingface.co/fixie-ai/ultravox-v0_7-glm-4_6"],"confidence":"high","unverified":"Language list, latency, compliance certifications, exact browser SDK method names (ultravox-client UltravoxSession.joinCall).","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Ultravox","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":5,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":true,"open_weights":true,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":true,"context_tokens":null,"byo_llm":null,"native_s2s":false,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.05,"_notes":"Audio-native understanding model (no ASR stage) with separate TTS output, so not a single speech-to-speech model. Open weights are the MIT audio adapter only. maxDuration defaults to 1 hour, no documented maximum. Concurrency hard cap 5 on PAYG; Pro has no hard cap. 30 free call minutes.","_added":["max_session_min (default 3600 s, no stated maximum, left null): https://docs.ultravox.ai/api-reference/calls/calls-post"]}},{"id":"phonic","name":"Phonic","vendor":"Phonic","category":"speech-to-speech","summary":"A voice-agent platform built around its own speech-to-speech model ('merritt'), with tools, evals, conversation replay, SIP and LiveKit support. Aimed at production phone agents where reliability matters more than lowest price.","status":"GA","models":[{"name":"merritt","status":"GA","notes":"The only STS model value in the API (default)."}],"transports":["WebSocket","Webhooks","SIP (Twilio, Telnyx)","LiveKit plugin","Amazon Connect via Chime SIP"],"audio":{"input":"pcm_44100 (default), pcm_24000, pcm_16000, pcm_8000, mulaw_8000","output":"Same options"},"languages":"51 languages (per docs index).","voices":"Named preset voices (e.g. sabrina, grant, eleanor, nolan) with accents; listable via voices.list.","latency":"Vendor claim: sub-500 ms speech-in to speech-out.","features":["webhook, WebSocket, context, transfer, MCP and built-in tools","interruption handling","pronunciation dictionary","multilingual switching","push-to-talk","conversation replay for testing","evals","audit logs and retention controls","background noise option"],"pricing":{"model":"per-minute","items":[{"what":"Conversation","price":"from $0.15","unit":"per minute","notes":"Usage-based 'starting at' price; no full rate card published."}],"audio_token_rate":"Not token-billed.","est_per_minute_usd":{"low":0.15,"high":0.15,"basis":"Published starting price; volume pricing via sales."},"free_tier":"Not stated.","source":"https://docs.phonic.ai/"},"limits":["Rate limits per organization: 500 requests/second shared across /sts/ws and SIP outbound; 5 requests/second for /conversations/outbound_call"],"regions":"Not stated.","setup":{"steps":["Create a Phonic account and API key; build an agent in the dashboard.","Connect with the SDK (or raw WebSocket to /v1/sts/ws) using Authorization: Bearer <key>; browsers use session tokens.","Send a config message naming the agent, then forward audio_chunk messages both ways."],"endpoint":"wss://api.phonic.ai/v1/sts/ws","auth":"Authorization: Bearer <PHONIC_API_KEY>; browser/mobile: ?session_token=... or your own JWTs","snippet_lang":"javascript","snippet":"// npm i phonic   (create an agent in the Phonic dashboard first)\nimport { PhonicClient } from \"phonic\";\n\nconst phonic = new PhonicClient({ apiKey: process.env.PHONIC_API_KEY });\n\n// Opens wss://api.phonic.ai/v1/sts/ws with Authorization: Bearer <key>\nconst socket = await phonic.conversations.connect();\n\nsocket.on(\"message\", (msg) => {\n  if (msg.type === \"audio_chunk\") play(msg.audio); // base64 agent audio (+ transcript)\n});\n\nawait socket.sendConfig({\n  type: \"config\",\n  agent: \"my-agent\",          // agent name from the dashboard\n  // input_format / output_format: pcm_44100 (default), pcm_24000, pcm_16000, pcm_8000, mulaw_8000\n});\n\n// Forward caller/mic audio (base64) as it arrives\nexport async function sendAudio(b64) {\n  await socket.sendAudioChunk({ type: \"audio_chunk\", audio: b64 });\n}\nfunction play(b64) {}"},"warnings":[{"severity":"medium","title":"Most expensive per minute here","detail":"From $0.15/min is 2-3x Ultravox or Deepgram Standard; check whether the price includes telephony and what volume discounts exist."},{"severity":"medium","title":"No public rate card","detail":"Only a 'starting at' price is published; enterprise terms vary."},{"severity":"medium","title":"Single proprietary model","detail":"Only 'merritt' is offered; you cannot swap in your own LLM for reasoning-heavy tasks."},{"severity":"low","title":"Default 44.1 kHz input","detail":"Default input format is pcm_44100; telephony bridges must set mulaw_8000 explicitly or audio will be wrong."},{"severity":"low","title":"Live conversation WebSocket is Preview","detail":"The docs label the live conversation WebSocket as Preview; expect protocol changes."}],"best_for":"Production phone agents in regulated settings (HIPAA) wanting a native speech-to-speech model with tooling for evals and replay.","open_source":false,"self_hostable":false,"compliance":"HIPAA and SOC 2 compliance, 99.9% uptime SLA (vendor claim).","docs":[{"label":"Docs","url":"https://docs.phonic.ai/"},{"label":"WebSocket conversations","url":"https://docs.phonic.ai/docs/build/with-web-sockets/via-web-sockets"},{"label":"API reference: conversations","url":"https://docs.phonic.ai/api-reference/conversations/conversations"}],"sources":["https://docs.phonic.ai/","https://docs.phonic.ai/llms.txt","https://docs.phonic.ai/api-reference/conversations/conversations.md","https://docs.phonic.ai/docs/platform/rate_limits.md","https://docs.phonic.ai/docs/build/agents/voices.md"],"confidence":"medium","unverified":"What the $0.15 covers (telephony?), concurrency caps, regions, free trial.","cat":"voice","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":500,"languages":51,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":true,"soc2":true,"gdpr_eu":null,"webrtc":false,"websocket":true,"sip":true,"open_weights":false,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":null,"context_tokens":null,"byo_llm":false,"native_s2s":true,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.15,"_notes":"Latency is a vendor claim of under 500 ms speech in to speech out. Price is the published starting rate. HIPAA and SOC 2 are vendor claims. LiveKit plugin available."}},{"id":"cartesia-agents","name":"Cartesia Managed Agents (Line)","vendor":"Cartesia","category":"voice-agent","summary":"Managed voice agents combining Cartesia Ink-2 speech recognition, an LLM you choose and Sonic-3.6 TTS, with phone numbers, SIP and evals. Suits teams that like Sonic's low-latency voices and want a managed agent runtime.","status":"GA","models":[{"name":"Ink-2 (STT) + chosen LLM + Sonic-3.6 (TTS)","status":"GA","notes":"Cascaded pipeline. LLM list per account via GET /v1/agents/models (e.g. gpt-5.4-mini)."}],"transports":["WebSocket","Phone numbers (Cartesia, Twilio import)","SIP trunking"],"audio":{"input":"pcm_16000, pcm_24000, pcm_44100 (16-bit) or mulaw_8000, headerless mono","output":"Streamed audio_output events (speaking-pace delivery option)"},"languages":"Configurable per agent; list not read.","voices":"Cartesia Sonic voice library.","latency":"Not stated on the pages read.","features":["system, webhook and client tools","dynamic variables","agent versioning","turn-taking and interruptions","outbound and batch calling","call metrics/evaluations"],"pricing":{"model":"per-minute","items":[{"what":"Agent calling","price":"$0.06","unit":"per minute","notes":"Base rate for all voice agent calls; excludes LLM."},{"what":"Cartesia-provided phone number telephony","price":"+$0.014","unit":"per minute","notes":""},{"what":"LLM","price":"per model token prices","unit":"per 1M tokens","notes":"Listed per account via GET /v1/agents/models. Free LLM usage promotion ran until October 1, 2026."}],"audio_token_rate":"Not token-billed (LLM portion is).","est_per_minute_usd":{"low":0.06,"high":0.09,"basis":"$0.06 base plus $0.014 telephony on Cartesia numbers plus LLM tokens (small models add little; own estimate)."},"free_tier":"Subscription plans include monthly credits; free LLM promo ended 2026-10-01.","source":"https://docs.cartesia.ai/pricing.md"},"limits":["Concurrency not documented on the pricing page","max_output_tokens per response 1-4096"],"regions":"Twilio regional routing documented; Cartesia hosting regions not stated.","setup":{"steps":["Create an agent in the Cartesia Playground or via API (instructions, voice, LLM, tools).","Server: open the agent WebSocket with X-API-Key and cartesia_version; browsers: mint an agent access token.","Send session_create with input_format, wait for session_ready, stream audio_input, play audio_output, clear on audio_output_clear."],"endpoint":"wss://api.cartesia.ai/v1/agents/websocket/{agent_id}?cartesia_version=2026-08-14","auth":"X-API-Key header (server) or access_token query parameter (browser/mobile)","snippet_lang":"javascript","snippet":"// npm i ws   (create the agent in the Cartesia Playground or via API first)\nimport WebSocket from \"ws\";\n\nconst ws = new WebSocket(\n  `wss://api.cartesia.ai/v1/agents/websocket/${process.env.CARTESIA_AGENT_ID}?cartesia_version=2026-08-14`,\n  { headers: { \"X-API-Key\": process.env.CARTESIA_API_KEY } } // browsers: ?access_token=... instead\n);\nlet ready = false;\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"session_create\",\n    audio: { input_format: \"pcm_16000\", output_delivery: \"speaking_pace\" },\n  }));\n});\n\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"session_ready\") { ready = true; console.log(\"call\", ev.call_id); }\n  if (ev.type === \"audio_output\") play(Buffer.from(ev.audio, \"base64\"));\n  if (ev.type === \"audio_output_clear\") stopPlayback(); // user interrupted\n});\n\n// headerless mono 16-bit PCM, 20-100 ms chunks\nexport function sendPcm(buf) {\n  if (ready) ws.send(JSON.stringify({ type: \"audio_input\", audio: buf.toString(\"base64\") }));\n}\nfunction play(pcm) {} function stopPlayback() {}"},"warnings":[{"severity":"high","title":"Free LLM period just ended","detail":"LLM usage was free until October 1, 2026; from now on LLM tokens are billed on top of $0.06/min. Budgets built during the promo are too low."},{"severity":"medium","title":"Versioned API header","detail":"Connections require cartesia_version (2026-08-14 in current docs); older version strings may behave differently."},{"severity":"medium","title":"Never ship the API key to clients","detail":"Browsers cannot set WebSocket headers; use short-lived agent access tokens."},{"severity":"medium","title":"Cascaded pipeline","detail":"STT, LLM and TTS are separate stages; tone and emotion in the caller's voice are not passed to the LLM."},{"severity":"low","title":"Telephony surcharge","detail":"Cartesia numbers add $0.014/min; Twilio import or SIP is billed by your carrier."}],"best_for":"Phone and web agents that want Sonic voices with a managed runtime and simple per-minute billing.","open_source":false,"self_hostable":false,"compliance":"Not stated on the pages read.","docs":[{"label":"Managed Agents","url":"https://docs.cartesia.ai/agents/introduction"},{"label":"Agent WebSocket API","url":"https://docs.cartesia.ai/line/integrations/websocket-api"},{"label":"LLMs","url":"https://docs.cartesia.ai/agents/models"},{"label":"Pricing","url":"https://docs.cartesia.ai/pricing"}],"sources":["https://docs.cartesia.ai/pricing.md","https://docs.cartesia.ai/agents/introduction.md","https://docs.cartesia.ai/line/integrations/websocket-api.md","https://docs.cartesia.ai/agents/models.md"],"confidence":"high","unverified":"Concurrency limits, languages, latency, compliance.","cat":"voice","kind":"voice","verified_at":"2026-10-10","short":"Cartesia Agents","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"sip":true,"open_weights":false,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":null,"context_tokens":null,"byo_llm":null,"native_s2s":false,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.06,"_notes":"$0.06/min base excludes LLM tokens; Cartesia phone numbers add $0.014/min. Subscription plans include monthly credits. LLM chosen from a per-account list."}},{"id":"assemblyai-voice-agent","name":"AssemblyAI Voice Agent API","vendor":"AssemblyAI","category":"voice-agent","summary":"A single WebSocket voice agent (STT, LLM via AssemblyAI's LLM Gateway, TTS, turn detection, tools) at an all-inclusive $4.50 per hour. Attractive when you want one flat price that already includes the LLM.","status":"GA","models":[{"name":"Universal-3.6 Pro realtime STT + LLM Gateway + TTS","status":"GA","notes":"Launched April 2026; formerly called the Speech-to-Speech API. Cascaded, runs on self-hosted LiveKit."}],"transports":["WebSocket"],"audio":{"input":"PCM16 mono 24 kHz (audio/pcm default)","output":"PCM16 24 kHz (reply.audio)"},"languages":"Declared via input.language_codes or auto-detected.","voices":"Preset voices (e.g. alba).","latency":"Not stated on the pages read.","features":["tool calling","turn detection (VAD thresholds, min/max silence)","interruptions","greeting","stored agents bound by agent_id","mid-session prompt updates","session resume via session_id"],"pricing":{"model":"per-minute","items":[{"what":"Voice Agent session","price":"$4.50","unit":"per hour ($0.075/min)","notes":"Vendor says this includes STT, LLM reasoning, TTS, turn detection, interruption handling and tool calling."}],"audio_token_rate":"Not token-billed.","est_per_minute_usd":{"low":0.075,"high":0.075,"basis":"Flat all-inclusive rate; telephony not included."},"free_tier":"$50 free credits at signup, no card.","source":"https://www.assemblyai.com/pricing.md"},"limits":["Session has a maximum duration (expires_at in session.ready); session_expired arrives with no advance warning","Streaming concurrency on the free tier: 5 new streams per minute (general streaming figure)"],"regions":"Not stated.","setup":{"steps":["Create an AssemblyAI account and API key.","Open wss://agents.assemblyai.com/v1/ws with Authorization: <key>.","Send session.update (either an agent_id OR inline system_prompt/greeting/voice, not both).","Wait for session.ready, then stream input.audio and play reply.audio."],"endpoint":"wss://agents.assemblyai.com/v1/ws","auth":"Authorization header with the raw API key (Bearer prefix also accepted)","snippet_lang":"javascript","snippet":"// npm i ws\nimport WebSocket from \"ws\";\n\nconst ws = new WebSocket(\"wss://agents.assemblyai.com/v1/ws\", {\n  headers: { Authorization: process.env.ASSEMBLYAI_API_KEY }, // raw key; \"Bearer <key>\" also accepted\n});\nlet ready = false;\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({\n    type: \"session.update\",\n    session: {\n      system_prompt: \"You are a friendly support agent. Keep responses under 2 sentences.\",\n      greeting: \"Hi! How can I help you today?\",\n      output: { voice: \"alba\" },\n    },\n  }));\n});\n\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"session.ready\") ready = true;          // only send audio after this\n  if (ev.type === \"reply.audio\") play(Buffer.from(ev.data, \"base64\")); // PCM16\n  if (ev.type === \"error\") console.error(ev);\n});\n\n// PCM16 mono 24 kHz\nexport function sendPcm(buf) {\n  if (ready) ws.send(JSON.stringify({ type: \"input.audio\", audio: buf.toString(\"base64\") }));\n}\nfunction play(pcm) {}"},"warnings":[{"severity":"medium","title":"Sessions expire without warning","detail":"session.ready includes expires_at; when the TTL is reached you get session_expired with no advance notice. Track the timestamp and reconnect with the session_id."},{"severity":"medium","title":"Docs disagree on the URL","detail":"The quickstart uses /v1/realtime while deploy and config docs use /v1/ws; use /v1/ws and confirm."},{"severity":"medium","title":"Telephony still maturing","detail":"The pricing page described Twilio SIP as 'coming Q2 2026' with price TBA; confirm current status before planning phone deployments."},{"severity":"low","title":"agent_id and inline config are exclusive","detail":"Sending system_prompt or tools alongside agent_id is rejected."},{"severity":"low","title":"Cascaded under the hood","detail":"Despite the old 'Speech-to-Speech API' name it chains STT, LLM and TTS."}],"best_for":"Teams wanting a simple all-in price with LLM included and strong STT accuracy.","open_source":false,"self_hostable":false,"compliance":"Vendor describes the API as PCI-certified; other certifications not re-verified.","docs":[{"label":"Voice Agent API","url":"https://assemblyai.com/docs/voice-agents/voice-agent-api"},{"label":"Session configuration","url":"https://assemblyai.com/docs/voice-agents/voice-agent-api/session-configuration"},{"label":"Events reference","url":"https://assemblyai.com/docs/voice-agents/voice-agent-api/events-reference"},{"label":"Pricing","url":"https://www.assemblyai.com/pricing"}],"sources":["https://www.assemblyai.com/pricing.md","https://assemblyai.com/docs/voice-agents/voice-agent-api/session-configuration","https://assemblyai.com/docs/voice-agents/voice-agent-api/events-reference","https://www.assemblyai.com/docs/voice-agents/voice-agent-api/deploy"],"confidence":"medium","unverified":"Which LLMs are included at the flat rate, max session length, concurrency, language list, current SIP status.","cat":"voice","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":50,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":null,"context_tokens":null,"byo_llm":null,"native_s2s":false,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":0.075,"_notes":"Flat $4.50/hour includes STT, LLM, TTS and tools. Sessions have an expiry (expires_at) whose length is not documented. Twilio SIP was listed as coming Q2 2026."}},{"id":"inworld-realtime","name":"Inworld Realtime API","vendor":"Inworld AI","category":"voice-agent","summary":"An OpenAI-Realtime-compatible voice API that chains Inworld STT, any of 100+ routed LLMs and Inworld TTS-2, over WebSocket or WebRTC. Good if you already speak the OpenAI Realtime protocol but want cheaper voices and free choice of LLM.","status":"GA","models":[{"name":"STT: inworld/inworld-stt-1","status":"GA","notes":"Default realtime STT with turn-taking controls."},{"name":"TTS: inworld-tts-2 (and TTS-2 Flash)","status":"GA","notes":""},{"name":"LLM: provider/model ids or routers (e.g. openai/gpt-4o-mini, inworld/latency-optimizer-ab-test)","status":"GA","notes":"Router offers 100+ models billed at provider cost."}],"transports":["WebSocket","WebRTC"],"audio":{"input":"PCM16 mono 24 kHz default; G.711 mu-law/A-law 8 kHz; float32","output":"Same options"},"languages":"Depends on STT/TTS models chosen; not listed on the pages read.","voices":"Inworld voice library (e.g. Clive) and cloned voices.","latency":"Not stated on the pages read.","features":["OpenAI Realtime-style events","semantic VAD with interrupt_response","function calling (tools, tool_choice)","mid-session model/voice switching","memory, back-channel and responsiveness extensions (providerData)"],"pricing":{"model":"per-character","items":[{"what":"Realtime TTS-2","price":"$25 (On-Demand) down to $12.50 (Growth)","unit":"per 1M characters","notes":"Enterprise as low as $5."},{"what":"TTS-2 Flash","price":"$15 down to $7","unit":"per 1M characters","notes":""},{"what":"STT 1","price":"$0.15 (On-Demand) / $0.10 (paid plans)","unit":"per hour","notes":""},{"what":"LLM","price":"provider cost","unit":"per token","notes":"Billed at cost, separately."},{"what":"Plans","price":"On-Demand free, Creator $25, Builder $100, Developer $300, Growth $1,500","unit":"per month","notes":"Paid plans include credits equal to the plan price."}],"audio_token_rate":"Inworld estimates 1 minute of speech is about 1,000 characters of TTS.","est_per_minute_usd":{"low":0.01,"high":0.03,"basis":"Own estimate: STT about $0.0025/min plus TTS-2 about $0.0125-0.025 per minute of agent speech (1,000 chars/min), before LLM cost."},"free_tier":"On-Demand plan free to start with up to 70 TTS minutes and up to 400 STT minutes.","source":"https://inworld.ai/pricing"},"limits":["Concurrent requests: 5 (On-Demand), 10 (Creator), 50 (Builder), 150 (Developer), 500 (Growth), custom (Enterprise)"],"regions":"Not stated.","setup":{"steps":["Create an Inworld account and API key in the Portal.","Server: open the realtime WebSocket with Basic auth; browser: mint a JWT session token and use Bearer auth (or WebRTC).","On session.created send session.update choosing LLM, STT, TTS model and voice; stream input_audio_buffer.append and play response.output_audio.delta."],"endpoint":"wss://api.inworld.ai/api/v1/realtime/session?key=<session-id>&protocol=realtime","auth":"Server: Authorization: Basic <api-key>; browser: Authorization: Bearer <jwt>","snippet_lang":"javascript","snippet":"// npm i ws   (server-side; browsers use a short-lived JWT with Bearer auth)\nimport WebSocket from \"ws\";\n\nconst sessionId = crypto.randomUUID();\nconst ws = new WebSocket(\n  `wss://api.inworld.ai/api/v1/realtime/session?key=${sessionId}&protocol=realtime`,\n  { headers: { Authorization: `Basic ${process.env.INWORLD_API_KEY}` } }\n);\n\nws.on(\"message\", (raw) => {\n  const ev = JSON.parse(raw.toString());\n  if (ev.type === \"session.created\") {\n    ws.send(JSON.stringify({\n      type: \"session.update\",\n      session: {\n        type: \"realtime\",\n        model: \"openai/gpt-4o-mini\",          // LLM billed at provider cost\n        instructions: \"You are a friendly narrator.\",\n        output_modalities: [\"audio\", \"text\"],\n        audio: {\n          input: { transcription: { model: \"inworld/inworld-stt-1\" },\n                   turn_detection: { type: \"semantic_vad\", create_response: true, interrupt_response: true } },\n          output: { voice: \"Clive\", model: \"inworld-tts-2\" },\n        },\n      },\n    }));\n  }\n  if (ev.type === \"response.output_audio.delta\") play(Buffer.from(ev.delta, \"base64\"));\n});\n\n// 24 kHz mono PCM16, 60-100 ms chunks (OpenAI-style event)\nexport function sendPcm(buf) {\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.append\", audio: buf.toString(\"base64\") }));\n}\nfunction play(pcm) {}"},"warnings":[{"severity":"high","title":"LLM cost is extra","detail":"Published prices cover STT and TTS only; the routed LLM is billed at provider cost on top."},{"severity":"medium","title":"OpenAI-compatible, not identical","detail":"Event names follow OpenAI Realtime, but Inworld options live in providerData and full drop-in compatibility is not claimed. Test your existing client."},{"severity":"medium","title":"Low concurrency on free tier","detail":"On-Demand allows 5 concurrent requests; production needs at least Builder (50)."},{"severity":"low","title":"Cascaded pipeline","detail":"STT, LLM and TTS are separate; vocal emotion is not passed to the LLM."},{"severity":"low","title":"Character-based TTS billing","detail":"Verbose LLM output directly increases TTS cost; cap max_output_tokens."}],"best_for":"Games, characters and consumer apps that want OpenAI-Realtime-style integration with cheaper expressive TTS and free LLM choice.","open_source":false,"self_hostable":false,"compliance":"Not stated on the pages read.","docs":[{"label":"Realtime WebSocket","url":"https://docs.inworld.ai/realtime/connect/websocket"},{"label":"Configuring models","url":"https://docs.inworld.ai/realtime/usage/using-realtime-models"},{"label":"Pricing","url":"https://inworld.ai/pricing"}],"sources":["https://docs.inworld.ai/realtime/connect/websocket.md","https://docs.inworld.ai/realtime/usage/using-realtime-models.md","https://inworld.ai/pricing"],"confidence":"medium","unverified":"Latency, languages, regions, compliance; whether realtime sessions carry any per-minute platform fee beyond STT/TTS/LLM.","cat":"voice","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":5,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":false,"open_weights":false,"function_calling":true,"image_input":null,"voices":null,"voice_cloning":true,"context_tokens":null,"byo_llm":null,"native_s2s":false,"price_audio_in_per_1m":null,"price_audio_out_per_1m":null,"price_per_min":null,"_notes":"Billed per TTS character and STT hour, not per minute; vendor-based estimate about $0.01 to $0.03 per minute before LLM. LLM chosen from a router of 100+ provider models, billed at cost; custom endpoint not stated. Concurrency 5 on On-Demand, 50 on Builder."}},{"id":"kyutai-moshi","name":"Moshi","vendor":"Kyutai","category":"open-model","summary":"The original open full-duplex speech-to-speech model: it listens and talks at the same time with about 200 ms latency. Great for research and natural-sounding chit-chat demos; too small and knowledge-light for most business agents.","status":"GA","models":[{"name":"kyutai/moshiko-pytorch-bf16 (male voice), kyutai/moshika-pytorch-bf16 (female voice)","status":"GA","notes":"About 7.7B params; also q8 PyTorch (experimental), MLX q4/q8/bf16 and Candle q8/bf16 variants."},{"name":"kyutai/moshika-rag-pytorch-bf16, kyutai/moshika-rl-seamless","status":"Preview","notes":"Newer 2026 research checkpoints on Hugging Face (RAG and RL fine-tunes); details not verified."}],"transports":["WebSocket (built-in web server and client)"],"audio":{"input":"24 kHz via the Mimi codec (12.5 Hz frames, 80 ms)","output":"24 kHz Mimi-decoded speech"},"languages":"English (not stated on the repo page; the released models were trained for English).","voices":"Two fixed voices: Moshiko (male) and Moshika (female).","latency":"Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU.","features":["full duplex (overlapping speech, backchannels)","barge-in","streaming Mimi codec","PyTorch, MLX (Apple Silicon) and Rust/Candle backends"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://github.com/kyutai-labs/moshi"},"limits":["Small 7B text backbone: limited knowledge and reasoning","No built-in tool/function calling","Context length not stated on the repo"],"regions":"Self-hosted anywhere.","setup":{"steps":["Linux with an NVIDIA GPU (24 GB for PyTorch) or a Mac with Apple Silicon (MLX).","pip install moshi (or moshi_mlx on Mac).","Start the server and open the web UI it serves."],"endpoint":"Local: https://localhost:8998 (web UI and WebSocket served by moshi.server)","auth":"None by default; put it behind your own auth/proxy.","snippet_lang":"python","snippet":"# NVIDIA GPU, PyTorch (about 24 GB VRAM)\npip install -U moshi\npython -m moshi.server --hf-repo kyutai/moshika-pytorch-bf16\n\n# Apple Silicon (MLX, 4-bit)\npip install -U moshi_mlx\npython -m moshi_mlx.local -q 4 --hf-repo kyutai/moshika-mlx-q4\n\n# Rust/Candle server (CUDA)\n# cargo run --features cuda --bin moshi-backend -r -- --config moshi-backend/config.json standalone"},"hardware":"PyTorch: NVIDIA GPU with about 24 GB VRAM. MLX q4/q8 tested on MacBook Pro M3. Rust backend needs CUDA or Metal.","license":"Weights CC-BY-4.0; Python code MIT; Rust backend Apache-2.0.","warnings":[{"severity":"high","title":"Not an agent brain","detail":"The 7B backbone has thin world knowledge and no function calling; it is a conversational demo model, not a support agent."},{"severity":"medium","title":"English only, two voices","detail":"No voice selection beyond the two fine-tunes and no multilingual support in the main release."},{"severity":"medium","title":"24 GB VRAM for PyTorch","detail":"PyTorch quantization is limited (q8 marked experimental); smaller GPUs need the MLX or Candle builds."},{"severity":"medium","title":"Attribution required","detail":"CC-BY-4.0 weights allow commercial use but require attribution to Kyutai."},{"severity":"low","title":"No auth or multi-tenant serving","detail":"The bundled server is single-user oriented; batching and auth are your job."},{"severity":"low","title":"Research cadence","detail":"New checkpoints (RAG, RL 'seamless') appear on Hugging Face without stable product docs."}],"best_for":"Research, latency benchmarks, and natural-feeling chit-chat demos on a single GPU or Mac.","open_source":true,"self_hostable":true,"compliance":"Self-hosted; compliance is your responsibility.","docs":[{"label":"GitHub","url":"https://github.com/kyutai-labs/moshi"},{"label":"Model (moshika)","url":"https://huggingface.co/kyutai/moshika-pytorch-bf16"}],"sources":["https://github.com/kyutai-labs/moshi","https://huggingface.co/api/models/kyutai/moshiko-pytorch-bf16","https://huggingface.co/api/models?author=kyutai"],"confidence":"high","unverified":"Exact MLX repo name in the snippet; context limit; details of the 2026 RAG/RL checkpoints.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":200,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":7.7,"vram_gb":24,"cpu_ok":false,"commercial_ok":true,"license_short":"CC-BY-4.0","full_duplex":true,"streaming":true,"_notes":"200 ms on an L4 GPU (160 ms theoretical). 24 GB VRAM for PyTorch; MLX quantized builds run on Apple Silicon. Attribution required."}},{"id":"kyutai-unmute","name":"Unmute","vendor":"Kyutai","category":"open-model","summary":"An open, low-latency cascaded voice stack: Kyutai streaming STT plus any OpenAI-compatible text LLM plus Kyutai streaming TTS. The practical way to give a smart self-hosted LLM a natural real-time voice.","status":"GA","models":[{"name":"kyutai/stt-1b-en_fr, kyutai/stt-2.6b-en","status":"GA","notes":"Streaming STT."},{"name":"kyutai/tts-1.6b-en_fr","status":"GA","notes":"Streaming TTS."},{"name":"Any OpenAI-compatible LLM (default Gemma 3 1B; vLLM, Ollama, OpenRouter)","status":"GA","notes":""}],"transports":["WebSocket (OpenAI-Realtime-like backend)","Web frontend"],"audio":{"input":"Browser mic via the bundled frontend","output":"Streamed TTS audio"},"languages":"English and French for STT/TTS (per model names); LLM language depends on your model.","voices":"Kyutai TTS voice set (kyutai/tts-voices).","latency":"Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs.","features":["plug in any OpenAI-compatible LLM","semantic voice activity / turn-taking","voice selection","Docker Compose and Swarm deployment"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint."},"free_tier":"Open weights","source":"https://github.com/kyutai-labs/unmute"},"limits":["No built-in tool calling (wrap your LLM server to add it)","Linux x86_64 only (Windows via WSL); no aarch64 or native Mac"],"regions":"Self-hosted.","setup":{"steps":["Linux x86_64 with a CUDA GPU (16 GB is enough for the default Gemma 3 1B setup).","Set your Hugging Face token.","Run docker compose; optionally point KYUTAI_LLM_URL at your own OpenAI-compatible LLM."],"endpoint":"Local web app served by the compose stack","auth":"None by default.","snippet_lang":"python","snippet":"git clone https://github.com/kyutai-labs/unmute && cd unmute\nexport HUGGING_FACE_HUB_TOKEN=hf_...\n# optional: use your own LLM (Ollama, vLLM, OpenRouter...)\n# export KYUTAI_LLM_URL=http://host.docker.internal:11434\n# export KYUTAI_LLM_MODEL=llama3.1\n# export KYUTAI_LLM_API_KEY=...\ndocker compose up --build"},"hardware":"CUDA GPU with at least 16 GB VRAM for the default setup; three or more GPUs recommended to split STT, TTS and LLM for lower latency.","license":"Unmute code MIT; Kyutai STT/TTS weights CC-BY-4.0; LLM licence depends on your choice.","warnings":[{"severity":"medium","title":"Cascaded, not native","detail":"Tone and emotion in the user's voice are not passed to the LLM; it trades that for a much smarter brain than Moshi."},{"severity":"medium","title":"No tool calling out of the box","detail":"Kyutai suggests wrapping vLLM in your own server to return tool-call results; Unmute itself is tool-unaware."},{"severity":"medium","title":"Platform restrictions","detail":"x86_64 Linux only (or WSL); no Mac or ARM servers."},{"severity":"low","title":"English/French speech only","detail":"STT/TTS models cover English and French."},{"severity":"low","title":"GPU count drives latency","detail":"Single-GPU deployments roughly double TTS latency versus split GPUs."}],"best_for":"Self-hosting a voice front-end for your own (possibly private) LLM with good latency.","open_source":true,"self_hostable":true,"compliance":"Self-hosted; your responsibility.","docs":[{"label":"GitHub","url":"https://github.com/kyutai-labs/unmute"}],"sources":["https://github.com/kyutai-labs/unmute","https://huggingface.co/api/models?author=kyutai"],"confidence":"high","unverified":"Exact environment variable values for Ollama in the snippet; full language list.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":750,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":null,"vram_gb":16,"cpu_ok":false,"commercial_ok":true,"license_short":"CC-BY-4.0","full_duplex":false,"streaming":true,"_notes":"Cascaded stack. 750 ms is TTS latency on one L40S, about 450 ms with services on separate GPUs. Code MIT, STT/TTS weights CC-BY-4.0; LLM licence depends on your choice. English and French speech."}},{"id":"nvidia-personaplex","name":"PersonaPlex-7B","vendor":"NVIDIA","category":"open-model","summary":"NVIDIA's full-duplex speech-to-speech model built on Moshi, adding a voice prompt and a text persona prompt so you can set who the agent is and how it sounds. The best open option for persona-driven, interruptible English voice agents.","status":"GA","models":[{"name":"nvidia/personaplex-7b-v1","status":"GA","notes":"About 8B params on disk; released January 2026; gated (accept licence)."},{"name":"kyutai/personaplex-rl-seamless","status":"Preview","notes":"June 2026 RL fine-tune published by Kyutai; not verified."}],"transports":["WebSocket (moshi.server web UI)"],"audio":{"input":"24 kHz","output":"24 kHz"},"languages":"English.","voices":"Set by an audio voice prompt (voice conditioning) plus a text persona prompt.","latency":"Vendor FullDuplexBench: smooth turn-taking latency 0.170, user-interruption latency 0.240 (units not given on the card).","features":["full duplex","barge-in and overlapping speech","persona text prompt","voice prompt conditioning"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://huggingface.co/nvidia/personaplex-7b-v1"},"limits":["English only","No documented tool calling (community forks exist)"],"regions":"Self-hosted.","setup":{"steps":["Accept the NVIDIA Open Model License on Hugging Face (gated).","Linux with an Ampere/Hopper GPU (A100 80 GB used in testing).","Run the Moshi server pointed at the PersonaPlex repo and open the local web UI."],"endpoint":"Local: https://localhost:8998","auth":"Hugging Face token for the gated download; server has no auth by default.","snippet_lang":"python","snippet":"pip install moshi\nhuggingface-cli login   # gated model: accept the licence first\npython -m moshi.server --hf-repo \"nvidia/personaplex-7b-v1\"\n# then open https://localhost:8998"},"hardware":"NVIDIA lists Ampere (A100) and Hopper (H100), tested on A100 80 GB. Community reports about 19 GB used on a 24 GB A10G; 24 GB is a safe target (third-party).","license":"NVIDIA Open Model License (model card says ready for commercial use); base Moshi weights CC-BY-4.0.","warnings":[{"severity":"medium","title":"Custom NVIDIA licence","detail":"Not Apache/MIT: read the NVIDIA Open Model License terms (and keep the CC-BY attribution for Moshi) before shipping."},{"severity":"medium","title":"Gated download","detail":"You must accept the licence and share contact details on Hugging Face; automate with a token."},{"severity":"medium","title":"VRAM unclear officially","detail":"NVIDIA only lists A100/H100; plan for 24 GB+ and test on smaller cards."},{"severity":"medium","title":"English only, no tools","detail":"Same Moshi limits: English speech, no built-in function calling."},{"severity":"low","title":"Latency claims are benchmark scores","detail":"The 0.170 / 0.240 figures are FullDuplexBench scores without stated units; measure on your hardware."}],"best_for":"Persona-driven English voice characters and interruptible assistants self-hosted on a single data-centre GPU.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"Model card","url":"https://huggingface.co/nvidia/personaplex-7b-v1"}],"sources":["https://huggingface.co/nvidia/personaplex-7b-v1","https://huggingface.co/api/models/nvidia/personaplex-7b-v1","https://themenonlab.blog/blog/nvidia-personaplex-full-duplex-voice-ai-how-to-guide","https://huggingface.co/abhinavpgagi/personaplex-tool-calling"],"confidence":"medium","unverified":"Minimum VRAM (third-party), latency units, details of Kyutai's RL variant.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":8,"vram_gb":24,"cpu_ok":false,"commercial_ok":true,"license_short":"Custom","full_duplex":true,"streaming":true,"_notes":"NVIDIA Open Model License (card says commercial use ok). FullDuplexBench latency scores 0.170 / 0.240 have no stated units. 24 GB VRAM is a third-party figure; NVIDIA lists A100/H100."}},{"id":"qwen3-omni-open","name":"Qwen3-Omni-30B-A3B (open weights)","vendor":"Alibaba Qwen","category":"open-model","summary":"Open-weight version of Qwen's omni model: hears speech, sees images/video and talks back, under Apache-2.0. Powerful and multilingual, but heavy to run and the standard vLLM server does not produce speech yet.","status":"GA","models":[{"name":"Qwen/Qwen3-Omni-30B-A3B-Instruct","status":"GA","notes":"MoE, about 35B total params, 3B active; speech output via the 'talker'."},{"name":"Qwen/Qwen3-Omni-30B-A3B-Thinking","status":"GA","notes":"Text output only (reasoning)."},{"name":"Qwen/Qwen2.5-Omni-7B / 3B","status":"GA","notes":"Older, smaller omni models (Apache-2.0) if 30B is too big."}],"transports":["Python (Transformers)","vLLM (text output only today)"],"audio":{"input":"Speech in 19 languages (18 listed)","output":"Speech in 10 languages"},"languages":"Text 119 languages; speech input about 19; speech output 10 (EN, ZH, FR, DE, RU, IT, ES, PT, JA, KO).","voices":"Ethan (default), Chelsie, Aiden.","latency":"Not given as a number on the model card.","features":["audio, image and video input","speech output with 3 voices","thinking variant"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. Hosted equivalents are on Alibaba Model Studio (see the Qwen-Omni Realtime entry)."},"free_tier":"Open weights","source":"https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct"},"limits":["vLLM serving supports only the thinker (no audio output) at the time of the card","No turn-key realtime/duplex server in the release"],"regions":"Self-hosted.","setup":{"steps":["Provision about 80 GB+ of GPU memory for BF16.","Install Transformers from source plus qwen-omni-utils (and flash-attn).","Load Qwen3OmniMoeForConditionalGeneration and call generate with speaker set; build your own streaming loop for realtime use."],"endpoint":"None (library)","auth":"n/a","snippet_lang":"python","snippet":"pip install git+https://github.com/huggingface/transformers accelerate\npip install -U qwen-omni-utils flash-attn --no-build-isolation\n\n# python\nfrom transformers import Qwen3OmniMoeForConditionalGeneration, Qwen3OmniMoeProcessor\nm = \"Qwen/Qwen3-Omni-30B-A3B-Instruct\"\nmodel = Qwen3OmniMoeForConditionalGeneration.from_pretrained(m, dtype=\"auto\", device_map=\"auto\", attn_implementation=\"flash_attention_2\")\nproc = Qwen3OmniMoeProcessor.from_pretrained(m)\n# build inputs with qwen_omni_utils.process_mm_info, then:\n# text_ids, audio = model.generate(**inputs, speaker=\"Ethan\")"},"hardware":"Vendor 'theoretical minimum' BF16 memory: 78.85 GB (15 s video) to 144.81 GB (120 s video) for Instruct; think 2x A100/H100 80 GB or one H200 class card. Audio-only use needs less but no official figure.","license":"Apache-2.0","warnings":[{"severity":"high","title":"No speech output from vLLM","detail":"The recommended vLLM path served only the text thinker at release; speech output needs the Transformers path (slow) or newer tooling. Check current vLLM support before planning."},{"severity":"high","title":"Big GPU bill","detail":"Roughly 80 GB+ VRAM in BF16 per the vendor table; not a single consumer GPU model."},{"severity":"medium","title":"Not full duplex","detail":"It is a turn-based omni model; you must build VAD, streaming and barge-in yourself."},{"severity":"medium","title":"Hosted versions are newer","detail":"Alibaba's hosted realtime models (qwen3.5/3.8-omni) are not open-weight; the open release is the 2025 Qwen3-Omni."},{"severity":"low","title":"Only 3 voices","detail":"Voice choice is limited to Ethan, Chelsie and Aiden."}],"best_for":"Self-hosted multilingual multimodal assistants where licence freedom (Apache-2.0) matters and data-centre GPUs are available.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"Model card","url":"https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct"}],"sources":["https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct","https://huggingface.co/api/models?author=Qwen&search=Omni"],"confidence":"high","unverified":"Current vLLM / vLLM-Omni audio-output support; audio-only VRAM; exact processor/generate call shape in the snippet.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":10,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":35,"vram_gb":79,"cpu_ok":false,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":false,"streaming":null,"_notes":"30B-A3B MoE, about 35B total and 3B active. 10 speech-output languages; 19 speech input, 119 text. 78.85 GB BF16 is the vendor theoretical minimum (with 15 s video). vLLM served text output only at release."}},{"id":"minicpm-o","name":"MiniCPM-o 4.5","vendor":"OpenBMB (ModelBest / Tsinghua)","category":"open-model","summary":"A 9B open omni model that does full-duplex speech (and video) streaming and runs on a single 12-24 GB GPU or a Mac via llama.cpp. The most practical open full-duplex model for English/Chinese with vision.","status":"GA","models":[{"name":"openbmb/MiniCPM-o-4_5","status":"GA","notes":"9B total (SigLip2 + Whisper-medium + CosyVoice2 + Qwen3-8B); int4 and GGUF builds available."},{"name":"openbmb/MiniCPM-o-2_6","status":"GA","notes":"Previous generation."}],"transports":["Python (Transformers)","vLLM","SGLang","llama.cpp-omni","Ollama"],"audio":{"input":"Streaming speech (and video frames)","output":"Streaming speech"},"languages":"Real-time speech conversation in English and Chinese; text in 30+ languages.","voices":"Configurable voices.","latency":"Vendor efficiency table: time to first token 0.6 s; decoding 154 tok/s (bf16) and 212 tok/s (int4).","features":["full-duplex speech streaming","full-duplex omni (see, listen, speak)","decides whether to speak at 1 Hz","stable long speech output (over 1 min)","quantized builds"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://huggingface.co/openbmb/MiniCPM-o-4_5"},"limits":["Half-duplex speech streaming marked under development","Speech conversation is bilingual (EN/ZH) only"],"regions":"Self-hosted.","setup":{"steps":["GPU with 12 GB+ (int4/llama.cpp) or about 28 GB for the PyTorch web demo.","Install the pinned Transformers and minicpmo-utils[all] for TTS/streaming.","Load with trust_remote_code and use the streaming demo, or serve via vLLM/llama.cpp-omni."],"endpoint":"Local (your server)","auth":"n/a","snippet_lang":"python","snippet":"pip install \"transformers==4.51.0\" accelerate \"torch>=2.3.0,<=2.8.0\" \"torchaudio<=2.8.0\" \"minicpmo-utils[all]>=1.0.5\"\n\n# python\nfrom transformers import AutoModel\nmodel = AutoModel.from_pretrained(\"openbmb/MiniCPM-o-4_5\", trust_remote_code=True, torch_dtype=\"auto\").eval().cuda()\n\n# or serve (text path) with vLLM:\n# vllm serve openbmb/MiniCPM-o-4_5 --trust-remote-code --max-num-batched-tokens 2048 --port 8000"},"hardware":"bf16 about 19 GB, int4 about 11 GB (vendor table). llama.cpp-omni full-duplex: NVIDIA 12 GB+ or Apple M4 Max 24 GB+. PyTorch web demo: 28 GB+.","license":"Apache-2.0 per the model card (older MiniCPM releases had extra registration terms; re-check the repo licence file).","warnings":[{"severity":"medium","title":"trust_remote_code","detail":"Loading runs model code from the repo; pin a revision and review it before production."},{"severity":"medium","title":"Pinned dependency versions","detail":"Requires transformers 4.51.0 and torch <= 2.8.0; conflicts with newer stacks are likely. Use a dedicated environment or container."},{"severity":"medium","title":"Bilingual speech only","detail":"Speech conversation is English and Chinese even though text covers 30+ languages."},{"severity":"low","title":"Licence history","detail":"Earlier MiniCPM models used a custom licence with commercial registration; the 4.5 card says Apache-2.0, but confirm the LICENSE file at the revision you ship."},{"severity":"low","title":"1 Hz speak decision","detail":"Full-duplex turn decisions happen once per second, which can feel slower than Moshi-class models."}],"best_for":"On-prem or edge full-duplex voice plus vision assistants on a single mid-range GPU or Mac.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"Model card","url":"https://huggingface.co/openbmb/MiniCPM-o-4_5"}],"sources":["https://huggingface.co/openbmb/MiniCPM-o-4_5","https://huggingface.co/api/models/openbmb/MiniCPM-o-4_5"],"confidence":"high","unverified":"Real-world end-to-end voice latency; whether vLLM serving includes speech output.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":600,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":9,"vram_gb":11,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":true,"streaming":true,"_notes":"600 ms is vendor time to first token. 11 GB is int4 (bf16 about 19 GB; llama.cpp full duplex needs 12 GB+ or an M4 Max). Speech in English and Chinese; text 30+ languages. Re-check the licence file at your revision."}},{"id":"step-audio-2-mini","name":"Step-Audio 2 mini","vendor":"StepFun","category":"open-model","summary":"An 8B open end-to-end audio model from StepFun for speech understanding and speech conversation, with paralinguistic understanding and tool-calling claims. A good open base for Chinese/English voice research; the stronger Step-Audio 2 and 2.5 are API-only.","status":"GA","models":[{"name":"stepfun-ai/Step-Audio-2-mini","status":"GA","notes":"8B, BF16."},{"name":"stepfun-ai/Step-Audio-2-mini-Think, -Base","status":"GA","notes":"Reasoning and base variants."},{"name":"stepfun-ai/Step-Audio-R1 / R1.1","status":"GA","notes":"Larger (about 33B) audio reasoning models; gated."}],"transports":["Python (Gradio web demo / scripts)"],"audio":{"input":"Speech audio","output":"Speech audio"},"languages":"English and Chinese tagged; ASR also benchmarked on Cantonese, Japanese, Arabic and Chinese dialects.","voices":"Not documented on the card.","latency":"Not stated; real-time streaming not stated on the card.","features":["speech-to-speech conversation","paralinguistic and non-vocal understanding","tool calling and multimodal RAG (claims; tool benchmark shown for full Step-Audio 2)"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://huggingface.co/stepfun-ai/Step-Audio-2-mini"},"limits":["No documented streaming/duplex server","No VRAM figure"],"regions":"Self-hosted.","setup":{"steps":["Python 3.10+, PyTorch 2.3+ with CUDA.","Install the listed dependencies, clone the repo, download weights.","Run the Gradio web demo or examples.py."],"endpoint":"Local Gradio demo","auth":"n/a","snippet_lang":"python","snippet":"git clone https://github.com/stepfun-ai/Step-Audio2 && cd Step-Audio2\npip install transformers==4.49.0 torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml gradio\nhuggingface-cli download stepfun-ai/Step-Audio-2-mini --local-dir Step-Audio-2-mini\npython web_demo.py   # or: python examples.py"},"hardware":"Not stated; 8B BF16 weights suggest a 24 GB-class GPU (own estimate).","license":"Apache-2.0","warnings":[{"severity":"medium","title":"Not a realtime server","detail":"The release is a turn-based model with a Gradio demo; streaming, VAD and barge-in are up to you."},{"severity":"medium","title":"Mini is weaker than the API model","detail":"Benchmarks show the mini below Step-Audio 2 (e.g. paralinguistic 80.00 vs 83.09); tool-calling results are reported for the full model, not mini."},{"severity":"low","title":"Old pinned Transformers","detail":"Requires transformers 4.49.0; isolate the environment."},{"severity":"low","title":"Chinese/English focus","detail":"Speech conversation quality outside EN/ZH is not documented."}],"best_for":"Research and prototyping of Chinese/English speech agents with an Apache-2.0 licence.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"Model card","url":"https://huggingface.co/stepfun-ai/Step-Audio-2-mini"}],"sources":["https://huggingface.co/stepfun-ai/Step-Audio-2-mini","https://huggingface.co/api/models?author=stepfun-ai&search=Audio"],"confidence":"medium","unverified":"VRAM needs, GitHub repo URL in the snippet, streaming support.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":8,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":false,"streaming":null,"_notes":"Turn-based release with a Gradio demo; no streaming server. English and Chinese."}},{"id":"kimi-audio","name":"Kimi-Audio-7B-Instruct","vendor":"Moonshot AI","category":"open-model","summary":"Moonshot's open audio foundation model (built on Qwen2.5-7B) for speech recognition, audio understanding and end-to-end speech conversation. Moonshot offers no hosted realtime voice API, so this is the only Kimi voice option.","status":"GA","models":[{"name":"moonshotai/Kimi-Audio-7B-Instruct","status":"GA","notes":"About 10B params BF16; released April 2025."}],"transports":["Python library","Docker image"],"audio":{"input":"Speech audio","output":"Speech at 24 kHz (output_type='both' returns text + audio)"},"languages":"English and Chinese.","voices":"Not documented.","latency":"Chunk-wise streaming detokenizer for low-latency generation; no figure.","features":["ASR","audio QA and captioning","emotion and sound-event recognition","end-to-end speech conversation"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://huggingface.co/moonshotai/Kimi-Audio-7B-Instruct"},"limits":["Not a full-duplex realtime server","No hosted Moonshot realtime API found"],"regions":"Self-hosted.","setup":{"steps":["CUDA GPU (24 GB-class suggested, own estimate).","pip install the Kimi-Audio package or pull the Docker image.","Load KimiAudio with load_detokenizer=True and call generate."],"endpoint":"None (library)","auth":"n/a","snippet_lang":"python","snippet":"pip install git+https://github.com/MoonshotAI/Kimi-Audio.git \"transformers<5\"\n# or: docker pull moonshotai/kimi-audio:v0.1\n\n# python\nfrom kimia_infer.api.kimia import KimiAudio\nmodel = KimiAudio(model_path=\"moonshotai/Kimi-Audio-7B-Instruct\", load_detokenizer=True)\n# wav, text = model.generate(messages, output_type=\"both\")"},"hardware":"Not stated on the card; about 10B BF16 params suggests 24 GB+ VRAM (own estimate).","license":"MIT (Qwen2.5-derived code Apache-2.0).","warnings":[{"severity":"medium","title":"No realtime serving","detail":"Turn-based library; you build streaming, VAD and interruption."},{"severity":"medium","title":"Aging release","detail":"Released April 2025 and little updated since; newer open models (MiniCPM-o 4.5, Qwen3-Omni) are generally preferred."},{"severity":"low","title":"EN/ZH only","detail":"Only English and Chinese are tagged."},{"severity":"low","title":"Moonshot has no voice API","detail":"Kimi's hosted API is text/multimodal chat only; do not expect a managed realtime endpoint."}],"best_for":"Research on audio understanding plus speech replies with a permissive MIT licence.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"Model card","url":"https://huggingface.co/moonshotai/Kimi-Audio-7B-Instruct"}],"sources":["https://huggingface.co/moonshotai/Kimi-Audio-7B-Instruct","https://huggingface.co/api/models/moonshotai/Kimi-Audio-7B-Instruct"],"confidence":"medium","unverified":"VRAM; exact import path and generate signature in the snippet.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":10,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"MIT","full_duplex":false,"streaming":true,"_notes":"Named 7B (Qwen2.5-7B base), about 10B params total. Chunk-wise streaming detokenizer. English and Chinese."}},{"id":"glm-4-voice","name":"GLM-4-Voice-9B","vendor":"Zhipu AI (zai-org)","category":"open-model","summary":"Zhipu's 2024 open end-to-end Chinese/English voice chat model (9B LLM plus speech tokenizer and decoder). Historically important and still usable, but older than the other options here.","status":"GA","models":[{"name":"zai-org/glm-4-voice-9b (also THUDM/glm-4-voice-9b)","status":"GA","notes":"LLM part; needs glm-4-voice-tokenizer and glm-4-voice-decoder."},{"name":"kyutai/glm-4-voice-of-reason-9b","status":"Preview","notes":"2026 Kyutai research derivative; not verified."}],"transports":["Local model server + Gradio web demo"],"audio":{"input":"Speech (12.5 tokens per second tokenizer)","output":"Speech via flow-matching decoder"},"languages":"Chinese and English.","voices":"Instruction-controlled style (emotion, speed, dialect) per repo; no voice list.","latency":"Vendor: generation can start after as few as 10 speech tokens; synthesis needs as few as 20 output tokens.","features":["end-to-end speech chat","streaming decoding","int4 option"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://github.com/zai-org/GLM-4-Voice"},"limits":["Released October 2024; not full duplex"],"regions":"Self-hosted.","setup":{"steps":["CUDA GPU (bf16, or int4 for less VRAM).","Start the model server, then the web demo with tokenizer and decoder paths."],"endpoint":"Local: http://127.0.0.1:8888 (web demo), model server on port 10000","auth":"n/a","snippet_lang":"python","snippet":"git clone https://github.com/zai-org/GLM-4-Voice && cd GLM-4-Voice\npip install -r requirements.txt\ngit clone https://huggingface.co/THUDM/glm-4-voice-decoder\npython model_server.py --host localhost --model-path THUDM/glm-4-voice-9b --port 10000 --dtype bfloat16 --device cuda:0\npython web_demo.py --tokenizer-path THUDM/glm-4-voice-tokenizer --model-path THUDM/glm-4-voice-9b --flow-path ./glm-4-voice-decoder"},"hardware":"Not stated; bf16 9B model plus decoder suggests 24 GB-class GPU, int4 less (own estimate).","license":"Code Apache-2.0; weights under the GLM-4 model licence (custom; read before commercial use).","warnings":[{"severity":"medium","title":"Custom weight licence","detail":"Weights follow the GLM-4 model agreement, not Apache; commercial terms must be checked."},{"severity":"medium","title":"Dated model","detail":"From October 2024; newer open models outperform it."},{"severity":"low","title":"Three components to wire","detail":"LLM, tokenizer and decoder are separate downloads and processes."},{"severity":"low","title":"EN/ZH only","detail":"Only Chinese and English speech."}],"best_for":"Chinese-first research baselines.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"GitHub","url":"https://github.com/zai-org/GLM-4-Voice"}],"sources":["https://github.com/zai-org/GLM-4-Voice","https://huggingface.co/zai-org/glm-4-voice-9b"],"confidence":"medium","unverified":"VRAM; GLM-4 licence commercial terms; requirements install step.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":9,"vram_gb":null,"cpu_ok":null,"commercial_ok":null,"license_short":"Custom","full_duplex":false,"streaming":true,"_notes":"Weights under the custom GLM-4 model licence (code Apache-2.0); check commercial terms. Streaming decoding can start after about 10 speech tokens."}},{"id":"liquid-lfm2-audio","name":"LFM2.5-Audio-1.5B","vendor":"Liquid AI","category":"open-model","summary":"A tiny (1.5B) end-to-end speech and text model that does interleaved speech-to-speech chat and runs on CPU via GGUF. Ideal for on-device or edge voice, as long as your company is under the $10M revenue licence threshold.","status":"GA","models":[{"name":"LiquidAI/LFM2.5-Audio-1.5B (+ GGUF, ONNX)","status":"GA","notes":"About 1.2B LM params, 1.5B total."},{"name":"LiquidAI/LFM2.5-Audio-1.5B-JP","status":"GA","notes":"Japanese variant (May 2026)."},{"name":"LiquidAI/LFM2-Audio-1.5B","status":"GA","notes":"Previous version."}],"transports":["Python package (liquid-audio)","llama.cpp GGUF","ONNX"],"audio":{"input":"Speech","output":"Speech (interleaved with text)"},"languages":"English (Japanese variant available).","voices":"Not documented.","latency":"Designed for low-latency real-time conversation; no figure.","features":["interleaved speech-to-speech chat","sequential ASR/TTS modes","CPU inference via GGUF"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B"},"limits":["Small model: limited knowledge/reasoning","English only (main model)"],"regions":"Self-hosted / on-device.","setup":{"steps":["Check the licence threshold first.","pip install liquid-audio and launch the demo, or use the GGUF build with llama.cpp on CPU."],"endpoint":"Local Gradio demo on port 7860","auth":"n/a","snippet_lang":"python","snippet":"pip install liquid-audio\npip install \"liquid-audio[demo]\"\nliquid-audio-demo   # serves on http://localhost:7860\n# CPU: use LiquidAI/LFM2.5-Audio-1.5B-GGUF with llama.cpp"},"hardware":"Runs on CPU via GGUF; a small GPU speeds it up (bfloat16, flash-attn optional).","license":"LFM Open License v1.0: commercial use only for entities under $10M annual revenue; above that, use is not licensed.","warnings":[{"severity":"high","title":"Revenue cap in the licence","detail":"Commercial use is licensed only if you (or your legal entity) have under $10,000,000 annual revenue. Larger companies need a separate deal with Liquid AI."},{"severity":"medium","title":"Tiny model limits","detail":"1.5B parameters means weak knowledge and reasoning; pair with retrieval or keep tasks narrow."},{"severity":"medium","title":"English only","detail":"Main model is English; a separate JP model exists."},{"severity":"low","title":"No full-duplex server","detail":"Interleaved generation is turn-based; barge-in handling is up to you."}],"best_for":"On-device or edge voice assistants for small companies, offline kiosks, prototypes.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"Model card","url":"https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B"},{"label":"Licence","url":"https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B/blob/main/LICENSE"}],"sources":["https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B","https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B/blob/main/LICENSE","https://huggingface.co/api/models?author=LiquidAI&search=Audio"],"confidence":"high","unverified":"Latency numbers; exact llama.cpp invocation.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":1.5,"vram_gb":null,"cpu_ok":true,"commercial_ok":false,"license_short":"Custom","full_duplex":false,"streaming":null,"_notes":"LFM Open License v1.0: commercial use only for entities under $10M annual revenue. English (separate Japanese variant). Runs on CPU via GGUF."}},{"id":"fun-audio-chat","name":"Fun-Audio-Chat-8B","vendor":"FunAudioLLM (Alibaba Tongyi Fun team, attribution from repo branding)","category":"open-model","summary":"An 8B open large audio-language model for low-latency voice chat with speech-to-speech inference, speech function calling and voice empathy, under Apache-2.0. A newer open alternative for Chinese/English voice agents that need tool use.","status":"GA","models":[{"name":"FunAudioLLM/Fun-Audio-Chat-8B","status":"GA","notes":"About 9.5B params; December 2025."}],"transports":["Python scripts (infer_s2s.py, infer_s2t.py)"],"audio":{"input":"Speech (5 Hz shared backbone, 25 Hz refined head)","output":"Speech"},"languages":"English and Chinese.","voices":"Not documented on the card.","latency":"No millisecond figure; vendor claims the 5 Hz frame rate cuts GPU hours by nearly 50%.","features":["speech-to-speech","speech function calling","spoken QA","speech instruction following","voice empathy"],"pricing":{"model":"free","items":[],"audio_token_rate":"n/a (self-hosted)","est_per_minute_usd":{"low":null,"high":null,"basis":"No licence fee; you pay for GPU time. "},"free_tier":"Open weights","source":"https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B"},"limits":["Full-duplex not stated","No serving stack beyond example scripts"],"regions":"Self-hosted.","setup":{"steps":["GPU with about 24 GB.","Clone the repo with submodules, install PyTorch 2.8 (cu128) and requirements, download weights.","Run the speech-to-speech example script."],"endpoint":"None (scripts)","auth":"n/a","snippet_lang":"python","snippet":"git clone --recurse-submodules https://github.com/FunAudioLLM/Fun-Audio-Chat && cd Fun-Audio-Chat\npip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128\npip install -r requirements.txt\nhf download FunAudioLLM/Fun-Audio-Chat-8B --local-dir ./pretrained_models/Fun-Audio-Chat-8B\nexport PYTHONPATH=`pwd`\npython examples/infer_s2s.py"},"hardware":"About 24 GB GPU memory for inference (vendor); 4x80 GB for training.","license":"Apache-2.0","warnings":[{"severity":"medium","title":"Example scripts, not a server","detail":"You must build streaming, VAD and an API around it."},{"severity":"medium","title":"Strict environment","detail":"Python 3.12, PyTorch 2.8.0 and ffmpeg are required; the card's generic Transformers snippet does not match the repo."},{"severity":"low","title":"EN/ZH only","detail":"Only English and Chinese."},{"severity":"low","title":"Young project","detail":"Low download counts and limited community tooling so far."}],"best_for":"Self-hosted Chinese/English voice agents that need spoken function calling under a permissive licence.","open_source":true,"self_hostable":true,"compliance":"Self-hosted.","docs":[{"label":"Model card","url":"https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B"}],"sources":["https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B"],"confidence":"medium","unverified":"Vendor attribution to Alibaba Tongyi (inferred from branding); latency; duplex behaviour.","cat":"open","kind":"voice","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":9.5,"vram_gb":24,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":null,"_notes":"Named 8B, about 9.5B params. Example scripts only, no server. Speech function calling. English and Chinese."}},{"id":"deepgram","name":"Deepgram Nova-3 and Flux (streaming)","vendor":"Deepgram","category":"stt","summary":"Low-cost WebSocket streaming STT. Nova-3 is the general model for captions, meetings and multilingual audio; Flux is a separate voice-agent model with built-in end-of-turn detection on its own /v2 endpoint.","status":"GA","models":[{"name":"nova-3 (nova-3-general)","status":"GA","notes":"General streaming and batch model. language=multi gives code-switching across 10 languages (en, es, fr, de, hi, ru, pt, ja, it, nl); many single-language codes also listed."},{"name":"nova-3-medical","status":"GA","notes":"Medical vocabulary variant; English plus the same 10-language multi mode."},{"name":"nova-3-pharma","status":"Listed in docs","notes":"English only, pharma vocabulary."},{"name":"flux-general-en","status":"GA","notes":"Conversational STT for voice agents with model-integrated end-of-turn. Only on wss://api.deepgram.com/v2/listen (v1 does not work)."},{"name":"flux-general-multi","status":"Listed in docs","notes":"Flux multilingual, 10 languages, language_hint supported."},{"name":"nova-2 / enhanced / base","status":"Legacy","notes":"Still available at old rates (Nova-2 streaming $0.35/hr per pricing FAQ). Use Nova-2 only for languages Nova-3 lacks."}],"transports":["WebSocket"],"audio":{"input":"Raw linear16, linear32, mulaw, alaw, opus, ogg-opus need encoding + sample_rate (8000, 16000, 24000, 44100, 48000 listed for Flux); containerized WAV/Ogg/WebM can be sent without those params. Flux strongly recommends 80 ms chunks (2560 bytes at 16 kHz linear16).","output":"n/a"},"languages":"Nova-3: multilingual code-switching mode for 10 languages plus a long list of single-language codes (see Models and Languages page). Flux: English, or 10 languages with flux-general-multi.","latency":"Vendor claim: Flux end-of-turn detection around 260 ms. No published latency figure for Nova-3 found in this pass.","features":["interim results (Nova-3)","endpointing and utterance end (Nova-3)","model-integrated end-of-turn with EagerEndOfTurn / TurnResumed (Flux)","diarization (paid add-on in streaming, Nova-3)","keyterm prompting (paid add-on)","redaction (paid add-on)","smart formatting (included)","word timestamps","Configure message to change Flux settings mid-stream","ForceEndTurn for bring-your-own turn detection"],"pricing":{"model":"per-minute","items":[{"what":"Nova-3 monolingual streaming","price":"$0.0048 / $0.0042","unit":"per minute (Pay As You Go / Growth)","notes":"Pre-recorded is cheaper: $0.0043 / $0.0036."},{"what":"Nova-3 multilingual streaming","price":"$0.0058 / $0.0050","unit":"per minute (PAYG / Growth)","notes":"Pre-recorded $0.0052 / $0.0043."},{"what":"Flux English streaming","price":"$0.0065 / $0.0057","unit":"per minute (PAYG / Growth)","notes":""},{"what":"Flux Multilingual streaming","price":"$0.0078 / $0.0068","unit":"per minute (PAYG / Growth)","notes":""},{"what":"Speaker diarization (streaming)","price":"+$0.0020 / +$0.0017","unit":"per minute","notes":"Included free on pre-recorded, charged on streaming."},{"what":"Keyterm prompting","price":"+$0.0013 / +$0.0012","unit":"per minute","notes":""},{"what":"Redaction","price":"+$0.0020 / +$0.0017","unit":"per minute","notes":""},{"what":"Entity detection","price":"+$0.0017","unit":"per minute","notes":""}],"est_per_minute_usd":{"low":0.0042,"high":0.0098,"basis":"Low = Nova-3 mono on Growth; high = Flux Multilingual PAYG ($0.0078) + diarization ($0.0020) where supported. Add-ons stack."},"free_tier":"$200 free credit for new accounts, does not expire until used (pricing page).","source":"https://deepgram.com/pricing"},"limits":["Concurrency: 150 WebSocket STT streams on Pay As You Go, 225 on Growth (pricing page).","Connection closes with NET-0001 if no audio or KeepAlive for 10 seconds; send KeepAlive (text frame) every 3-5 s.","Whisper Cloud limited to 5 concurrent requests."],"regions":"Hosted API default region; EU endpoint and self-hosted deployment exist (not re-verified in this pass).","setup":{"steps":["Create an API key in the Deepgram console.","Open wss://api.deepgram.com/v1/listen with model=nova-3 (or wss://api.deepgram.com/v2/listen?model=flux-general-en for Flux).","Send header Authorization: Token <key>; for browsers mint a short-lived JWT server-side (token-based auth, 30 s TTL).","Stream raw audio frames, send KeepAlive during silence, send {\"type\":\"CloseStream\"} to flush and close."],"endpoint":"wss://api.deepgram.com/v1/listen (Nova-3) ; wss://api.deepgram.com/v2/listen (Flux)","auth":"Authorization: Token <API_KEY> header. Browser: temporary JWT from the token-based auth endpoint (30 second TTL to open the socket).","snippet_lang":"python","snippet":"import asyncio, json, os, websockets  # pip install websockets>=14\n\nURL = (\"wss://api.deepgram.com/v1/listen?model=nova-3&encoding=linear16\"\n       \"&sample_rate=16000&interim_results=true&endpointing=300\")\n\nasync def main():\n    hdr = {\"Authorization\": f\"Token {os.environ['DEEPGRAM_API_KEY']}\"}\n    async with websockets.connect(URL, additional_headers=hdr) as ws:\n        async def send():\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:  # 16-bit PCM mono\n                while chunk := f.read(3200):  # 100 ms\n                    await ws.send(chunk)\n                    await asyncio.sleep(0.1)  # pace at real time\n            await ws.send(json.dumps({\"type\": \"CloseStream\"}))\n        asyncio.create_task(send())\n        async for msg in ws:\n            d = json.loads(msg)\n            if d.get(\"type\") == \"Results\":\n                text = d[\"channel\"][\"alternatives\"][0][\"transcript\"]\n                if text:\n                    print(\"FINAL\" if d[\"is_final\"] else \"partial\", text)\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Flux needs /v2, not /v1","detail":"Flux only works on wss://api.deepgram.com/v2/listen with model=flux-general-en or flux-general-multi. model=flux is invalid and the v1 endpoint will not accept it. Its message types (TurnInfo, EndOfTurn) differ from Nova-3 Results messages, so it is not a drop-in swap."},{"severity":"high","title":"10-second idle timeout","detail":"If neither audio nor a KeepAlive text frame arrives for 10 s the socket closes with NET-0001. Send KeepAlive every 3-5 s during mute or hold; send it as a text frame, not binary."},{"severity":"medium","title":"Diarization is extra on streaming only","detail":"Speaker diarization costs +$0.0020/min PAYG on streaming but is included on pre-recorded. Flux's feature matrix does not list diarization at all."},{"severity":"medium","title":"Streaming prices are marked promotional","detail":"The pricing page shows the streaming rates with strikethrough figures labelled limited-time promotional. Budget for the possibility that the struck-through (higher) rates return."},{"severity":"medium","title":"Eager end-of-turn multiplies LLM calls","detail":"Setting eager_eot_threshold enables EagerEndOfTurn/TurnResumed; Deepgram's own docs warn this can raise LLM calls by 50-70% because some speculative turns get cancelled."},{"severity":"low","title":"Add-ons stack per minute","detail":"Keyterm prompting, redaction and entity detection each add their own per-minute charge on top of the base model rate."}],"best_for":"Cheap, high-concurrency English or multilingual live captions (Nova-3) and voice agents that want turn detection from the STT itself (Flux).","open_source":false,"self_hostable":true,"compliance":"Vendor markets HIPAA (BAA) and SOC 2 and offers self-hosted/on-prem deployment; not re-verified in this pass, confirm with sales.","docs":[{"label":"Pricing","url":"https://deepgram.com/pricing"},{"label":"Flux quickstart","url":"https://developers.deepgram.com/docs/flux/quickstart"},{"label":"Models and languages","url":"https://developers.deepgram.com/docs/models-languages-overview"},{"label":"Audio keep alive","url":"https://developers.deepgram.com/docs/audio-keep-alive"},{"label":"Token-based auth","url":"https://developers.deepgram.com/guides/fundamentals/token-based-authentication"}],"sources":["https://deepgram.com/pricing","https://developers.deepgram.com/docs/flux/quickstart","https://developers.deepgram.com/docs/flux/feature-overview.md","https://developers.deepgram.com/docs/audio-keep-alive.md","https://developers.deepgram.com/docs/models-languages-overview.md"],"confidence":"high","unverified":"Compliance claims, regional endpoints and Nova-3 latency numbers were not re-checked. Whether streaming billing counts wall-clock connection time or audio sent was not confirmed on a primary page.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"Deepgram","facts":{"cat":"stt","latency_ms":260,"languages":59,"max_session_min":null,"concurrency":150,"free_tier":true,"free_credit_usd":200,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_hour":0.288,"diarization_streaming":true,"turn_detection":true,"keyterms":true,"pii_redaction":true,"interim_results":true,"code_switching":true,"telephony_8k":true,"bills_silence":null,"_notes":"Price is Nova-3 monolingual PAYG ($0.0048/min, marked promotional); Flux English is $0.39/hr. Latency is Flux end-of-turn (~260 ms); no Nova-3 figure. 59 languages for Nova-3; code-switching across 10. Diarization, keyterms and redaction are paid add-ons. HIPAA, SOC 2 and EU endpoint are vendor claims, not re-verified. 10 s idle timeout, no session cap stated.","_added":["languages: https://developers.deepgram.com/docs/models-languages-overview"]}},{"id":"assemblyai","name":"AssemblyAI Universal-3.6 Pro Realtime and Universal-Streaming","vendor":"AssemblyAI","category":"stt","summary":"WebSocket streaming STT (API v3) with built-in turn detection. Universal-3.6 Pro is the default high-accuracy multilingual streaming model; Universal-Streaming English/Multilingual are the cheaper tier.","status":"GA","models":[{"name":"universal-3-6-pro","status":"GA (default)","notes":"Default if speech_model is omitted. 32 languages with native code-switching. Keyterms included, prompting (beta) and Voice Focus as paid add-ons."},{"name":"universal-3-5-pro","status":"Listed in docs","notes":"19 languages."},{"name":"universal-streaming-english","status":"GA","notes":"English only, cheapest tier."},{"name":"universal-streaming-multilingual","status":"GA","notes":"English, Spanish, German, French, Portuguese, Italian."},{"name":"u3-rt-pro / u3-pro","status":"Removed from current API reference","notes":"Older Universal-3 streaming IDs per pricing page note; migrate to universal-3-6-pro."}],"transports":["WebSocket"],"audio":{"input":"Mono 16-bit PCM by default with sample_rate set to the source; mulaw also documented historically; AAC (ADTS, encoding=aac) and Opus (encoding=ogg_opus or opus) accepted. Chunks must be 50-1000 ms and must not arrive faster than real time.","output":"n/a"},"languages":"Universal-3.6 Pro: 32 languages with code-switching; 3.5 Pro: 19; Universal-Streaming Multilingual: 6; English model: English only.","latency":"Vendor docs rate the Pro models 'Fastest' and Universal-Streaming 'Fast'; no numeric figure on the model page.","features":["partial and final turns (Turn messages with end_of_turn)","turn detection with min_turn_silence / max_turn_silence / vad_threshold and latency modes","keyterm prompting (up to 100 terms)","general prompting (beta, Pro)","streaming diarization and multichannel (paid)","medical mode (paid)","Voice Focus noise isolation (paid, Pro)","PII redaction and profanity filter pages for streaming","UpdateConfiguration mid-stream","session-end webhooks","US/EU data-zone endpoints"],"pricing":{"model":"per-hour","items":[{"what":"Universal-3.6 Pro Realtime","price":"$0.45","unit":"per hour of session","notes":"Billed on WebSocket open time, not audio sent."},{"what":"Universal-Streaming English","price":"$0.15","unit":"per hour of session","notes":"Keyterms +$0.04/hr on this model."},{"what":"Universal-Streaming Multilingual","price":"$0.15","unit":"per hour of session","notes":""},{"what":"Streaming speaker diarization","price":"+$0.12","unit":"per hour","notes":"Applies to the whole session."},{"what":"Medical Mode","price":"+$0.15","unit":"per hour","notes":""},{"what":"General prompting (beta)","price":"+$0.05","unit":"per hour","notes":"Universal-3.6 Pro only."},{"what":"Voice Focus","price":"+$0.10","unit":"per hour","notes":"Universal-3.6 Pro only."}],"est_per_minute_usd":{"low":0.0025,"high":0.0095,"basis":"$0.15/hr Universal-Streaming = $0.0025/min; $0.45/hr U3.6 Pro + $0.12/hr diarization = $0.0095/min. Session time, so idle open sockets cost the same."},"free_tier":"$50 free credit at signup, no card required.","source":"https://www.assemblyai.com/pricing"},"limits":["Rate limit is on NEW sessions per minute, not total open sessions: Free 5/min, paid 100+/min, auto-scales +10% when you use 70%+ of it.","Max session 3 hours (error 3008); sessions not terminated auto-close after 3 hours and are billed for the full time.","Audio chunk must be 50-1000 ms (error 3007); sending faster than real time is rejected (3007).","Too many sessions gives close code 1008/3009."],"regions":"Edge routing across AWS US and EU regions by default; data-zone endpoints keep data in US (streaming.us.assemblyai.com) or EU (streaming.eu.assemblyai.com). Self-hosted streaming offered.","setup":{"steps":["Get an API key from the dashboard.","Open wss://streaming.assemblyai.com/v3/ws?speech_model=universal-3-6-pro&sample_rate=16000 with header Authorization: <API_KEY>.","Stream 50-1000 ms binary chunks at real-time pace; read Begin, Turn and Termination messages.","Send {\"type\":\"Terminate\"} when the call ends, otherwise you keep paying.","For browsers, mint a temporary token server-side (expires_in_seconds, optional max_session_duration_seconds) and pass it as token= in the URL."],"endpoint":"wss://streaming.assemblyai.com/v3/ws (US: wss://streaming.us.assemblyai.com/v3/ws, EU: wss://streaming.eu.assemblyai.com/v3/ws)","auth":"Authorization header with the raw API key; browser clients use a temporary streaming token generated server-side.","snippet_lang":"python","snippet":"import asyncio, json, os, websockets  # pip install websockets>=14\nfrom urllib.parse import urlencode\n\nparams = {\"speech_model\": \"universal-3-6-pro\", \"sample_rate\": 16000}\nURL = \"wss://streaming.assemblyai.com/v3/ws?\" + urlencode(params)\n\nasync def main():\n    hdr = {\"Authorization\": os.environ[\"ASSEMBLYAI_API_KEY\"]}\n    async with websockets.connect(URL, additional_headers=hdr) as ws:\n        async def send():\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:  # 16-bit PCM mono\n                while chunk := f.read(3200):  # 100 ms (must be 50-1000 ms)\n                    await ws.send(chunk)\n                    await asyncio.sleep(0.1)  # never faster than real time\n            await ws.send(json.dumps({\"type\": \"Terminate\"}))  # stops billing\n        asyncio.create_task(send())\n        async for msg in ws:\n            d = json.loads(msg)\n            if d[\"type\"] == \"Turn\":\n                print(\"FINAL\" if d.get(\"end_of_turn\") else \"partial\", d[\"transcript\"])\n            elif d[\"type\"] == \"Termination\":\n                break\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Billed for socket open time","detail":"Streaming is billed per session duration, the time the WebSocket is open, not the audio you send. Idle time counts and add-ons like diarization apply to the whole session. Always send Terminate when the call ends."},{"severity":"high","title":"Forgotten sessions bill for 3 hours","detail":"Sessions that are never terminated auto-close after 3 hours and are billed for the full duration; they also keep counting against your new-session rate limit."},{"severity":"medium","title":"Strict chunk size and pacing","detail":"Chunks under 50 ms or over 1000 ms close the session with 3007, and so does sending audio faster than real time. Streaming a file needs a sleep between chunks."},{"severity":"medium","title":"Model IDs changed in 2026","detail":"Older u3-rt-pro / u3-pro streaming IDs are no longer in the API reference, and the default model is now universal-3-6-pro (more expensive than Universal-Streaming). Pin speech_model explicitly so cost does not change under you."},{"severity":"low","title":"Guardrail add-ons not confirmed for streaming pricing","detail":"The pricing page lists PII redaction/profanity/content moderation prices only for batch models, although streaming docs have PII and profanity pages. Confirm cost before enabling."}],"best_for":"Voice agents and live captioning that want turn detection built in, multilingual code-switching, and US/EU data residency on a simple hourly price.","open_source":false,"self_hostable":true,"compliance":"US and EU data-zone streaming endpoints documented. Vendor markets SOC 2 and HIPAA BAA; not re-verified in this pass.","docs":[{"label":"Pricing","url":"https://www.assemblyai.com/pricing"},{"label":"Model selection","url":"https://www.assemblyai.com/docs/streaming/select-the-speech-model"},{"label":"Endpoints and data zones","url":"https://www.assemblyai.com/docs/streaming/endpoints-and-data-zones"},{"label":"Rate limits","url":"https://www.assemblyai.com/docs/streaming/rate-limits"},{"label":"Errors and closures","url":"https://www.assemblyai.com/docs/streaming/common-session-errors-and-closures"},{"label":"Turn detection","url":"https://www.assemblyai.com/docs/streaming/turn-detection"}],"sources":["https://www.assemblyai.com/pricing","https://www.assemblyai.com/docs/streaming/universal-streaming","https://www.assemblyai.com/docs/streaming/endpoints-and-data-zones.md","https://www.assemblyai.com/docs/streaming/rate-limits.md","https://www.assemblyai.com/docs/streaming/common-session-errors-and-closures.md","https://www.assemblyai.com/docs/streaming/turn-detection.md"],"confidence":"high","unverified":"Numeric latency, compliance certifications, and exact temporary-token endpoint path (docs show GET generate-streaming-token reference) not re-checked.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"AssemblyAI","facts":{"cat":"stt","latency_ms":null,"languages":32,"max_session_min":180,"concurrency":null,"free_tier":true,"free_credit_usd":50,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_hour":0.45,"diarization_streaming":true,"turn_detection":true,"keyterms":true,"pii_redaction":true,"interim_results":true,"code_switching":true,"telephony_8k":null,"bills_silence":true,"_notes":"Price and languages are for the default Universal-3.6 Pro model; Universal-Streaming is $0.15/hr (English, or 6 languages). Limit is 100+ new sessions per minute on paid, not a concurrent-stream cap. Diarization +$0.12/hr. Streaming PII pricing unconfirmed. HIPAA and SOC 2 are vendor claims, not re-verified."}},{"id":"speechmatics","name":"Speechmatics Realtime and Agent STT","vendor":"Speechmatics","category":"stt","summary":"Long-running WebSocket realtime transcription with Enhanced/Standard models, realtime diarization and big custom dictionaries; a separate Agent STT endpoint (Linden 1) gives turn-based, speaker-labelled output for voice agents.","status":"GA","models":[{"name":"enhanced","status":"GA","notes":"Highest accuracy, single-language, medical variant available. Realtime and batch."},{"name":"standard","status":"GA (default)","notes":"Faster/cheaper; used if model is not set."},{"name":"linden-1","status":"GA (Agent STT only)","notes":"Voice-agent model on wss://.../v2/agent with turn detection, segmentation and diarization."},{"name":"melia-1","status":"Preview for realtime","notes":"Multilingual code-switching (56 languages in realtime preview page); batch is the documented mode."},{"name":"oak-1","status":"Batch only","notes":"Multilingual healthcare model."}],"transports":["WebSocket"],"audio":{"input":"raw pcm_f32le, pcm_s16le or mulaw with explicit sample_rate; or 'file' type for wav, mp3, aac, ogg, mpeg, amr, m4a, mp4, flac.","output":"n/a"},"languages":"Language packs per session (vendor markets 50+ languages; count not re-verified). Enhanced/Standard need a selected language; Melia 1 handles code-switching.","latency":"Configurable via max_delay (docs examples use 0.7 s); no other vendor latency figure captured.","features":["partials (enable_partials)","max_delay latency control","realtime diarization (speaker and channel)","speaker identification","custom dictionary up to 20,000 items","turn detection / force end of utterance","translation add-on","Agent STT with turn events and speaker-labelled segments"],"pricing":{"model":"per-hour","items":[{"what":"Realtime Standard","price":"$0.45","unit":"per hour","notes":"$0.30/hr with the model-training discount (-33%)."},{"what":"Realtime Enhanced","price":"$0.80","unit":"per hour","notes":"$0.54/hr with model-training discount."},{"what":"Linden 1 (Agent STT)","price":"$0.30","unit":"per hour","notes":"$0.20/hr with model-training discount."},{"what":"Melia 1","price":"$0.40","unit":"per hour","notes":"Listed price; realtime is preview."},{"what":"Translation add-on","price":"$0.65","unit":"per hour","notes":"Supported for realtime."}],"est_per_minute_usd":{"low":0.0033,"high":0.0133,"basis":"Linden 1 with training discount ($0.20/hr) up to Enhanced list ($0.80/hr). Subscriptions/credit packs cut up to 25%/20%."},"free_tier":"$100 in credits at sign-up, no card (pricing page).","source":"https://www.speechmatics.com/pricing"},"limits":["Concurrent realtime sessions: Free 2, Pro 50, Enterprise custom.","Session ends at 48 hours, after 1 hour without AddAudio, or after 3 minutes with no audio or ping/pong.","Custom dictionary over 20,000 items closes the socket with protocol_error."],"regions":"EU, US, AUS listed for models; global endpoint routes to nearest region; regional endpoints for residency.","setup":{"steps":["Create an API key in the Speechmatics portal.","Connect to wss://eu.rt.speechmatics.com/v2 (or global.rt.speechmatics.com/v2) with Authorization: Bearer <key>, or a short-lived JWT for browsers.","Send StartRecognition JSON with audio_format and transcription_config, wait for RecognitionStarted.","Send binary audio (AddAudio), read AddPartialTranscript / AddTranscript, finish with EndOfStream {last_seq_no}.","For voice agents use wss://global.rt.speechmatics.com/v2/agent (Agent STT)."],"endpoint":"wss://eu.rt.speechmatics.com/v2 ; wss://global.rt.speechmatics.com/v2 ; Agent STT: wss://global.rt.speechmatics.com/v2/agent","auth":"Authorization: Bearer <API_KEY>; browser clients use a temporary JWT (createSpeechmaticsJWT with type rt and a ttl), passed as ?jwt=.","snippet_lang":"python","snippet":"import asyncio, json, os, websockets  # pip install websockets>=14\n\nURL = \"wss://eu.rt.speechmatics.com/v2\"\nSTART = {\"message\": \"StartRecognition\",\n         \"audio_format\": {\"type\": \"raw\", \"encoding\": \"pcm_s16le\", \"sample_rate\": 16000},\n         \"transcription_config\": {\"language\": \"en\", \"model\": \"enhanced\",\n                                  \"enable_partials\": True, \"max_delay\": 0.7}}\n\nasync def main():\n    hdr = {\"Authorization\": f\"Bearer {os.environ['SPEECHMATICS_API_KEY']}\"}\n    async with websockets.connect(URL, additional_headers=hdr) as ws:\n        await ws.send(json.dumps(START))\n        async def send():\n            n = 0\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:\n                while chunk := f.read(3200):  # 100 ms\n                    await ws.send(chunk); n += 1\n                    await asyncio.sleep(0.1)\n            await ws.send(json.dumps({\"message\": \"EndOfStream\", \"last_seq_no\": n}))\n        async for msg in ws:\n            d = json.loads(msg)\n            if d[\"message\"] == \"RecognitionStarted\":\n                asyncio.create_task(send())\n            elif d[\"message\"] in (\"AddPartialTranscript\", \"AddTranscript\"):\n                kind = \"partial\" if d[\"message\"] == \"AddPartialTranscript\" else \"FINAL\"\n                print(kind, d[\"metadata\"][\"transcript\"])\n            elif d[\"message\"] in (\"EndOfTranscript\", \"Error\"):\n                print(d); break\n\nasyncio.run(main())"},"warnings":[{"severity":"medium","title":"3-minute silence kill","detail":"A session with no audio and no ping/pong for 3 minutes is ended (and after 1 hour without AddAudio). Keep sending audio or WebSocket pings during holds."},{"severity":"medium","title":"Free tier is 2 concurrent sessions","detail":"Free accounts can only run 2 realtime sessions at once; Pro is 50. Load tests on a free key fail early."},{"severity":"medium","title":"Multilingual models are not realtime GA","detail":"Melia 1 and Oak 1 are documented as batch; realtime Melia 1 is a preview page. Realtime GA is Enhanced/Standard, which need you to pick a language."},{"severity":"low","title":"Training discount means data use","detail":"The cheaper '-33%' prices come from opting into model training. Check that this is acceptable for your data before picking those rates."},{"severity":"low","title":"Agent STT is a different API","detail":"Linden 1 runs only on the /v2/agent endpoint with its own message schema and is SaaS only; it is not selectable with model= on the normal realtime API."}],"best_for":"Long sessions (up to 48 h), broadcast and meeting captions with realtime diarization and large custom vocabularies; EU-hosted processing.","open_source":false,"self_hostable":true,"compliance":"Docs state Realtime SaaS does not store audio, transcripts or configuration. On-prem containers and Kubernetes deployments documented. Certifications not re-verified.","docs":[{"label":"Pricing","url":"https://www.speechmatics.com/pricing"},{"label":"Models","url":"https://docs.speechmatics.com/speech-to-text/models"},{"label":"Realtime limits","url":"https://docs.speechmatics.com/speech-to-text/realtime/limits"},{"label":"Realtime input formats","url":"https://docs.speechmatics.com/speech-to-text/realtime/input"},{"label":"Realtime API reference","url":"https://docs.speechmatics.com/api-ref/realtime-transcription-websocket"},{"label":"Agent STT","url":"https://docs.speechmatics.com/speech-to-text/agent-stt"}],"sources":["https://www.speechmatics.com/pricing","https://docs.speechmatics.com/speech-to-text/models.md","https://docs.speechmatics.com/speech-to-text/realtime/limits.md","https://docs.speechmatics.com/speech-to-text/realtime/input.md","https://docs.speechmatics.com/speech-to-text/agent-stt.md"],"confidence":"high","unverified":"Exact language count and whether realtime is billed by session length or audio length were not confirmed.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":50,"max_session_min":2880,"concurrency":50,"free_tier":true,"free_credit_usd":100,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_hour":0.45,"diarization_streaming":true,"turn_detection":true,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":true,"telephony_8k":true,"bills_silence":null,"_notes":"Languages: vendor markets 50+, count not verified. Price is Realtime Standard (default); Enhanced $0.80/hr, Linden 1 agent model $0.30/hr. Concurrency is the Pro plan (Free 2). Session max 48 h. Realtime code-switching only via Melia 1, which is a preview. Latency tunable via max_delay (0.7 s in examples), no vendor claim."}},{"id":"gladia","name":"Gladia Live (Solaria-1)","vendor":"Gladia","category":"stt","summary":"Two-step live API: POST to create a session, then stream audio to the returned WebSocket URL. Live uses Solaria-1 with 100+ languages and code-switching.","status":"GA","models":[{"name":"solaria-1","status":"GA (default, only live model)","notes":"100+ languages, code-switching, async and live."},{"name":"solaria-3","status":"GA (pre-recorded only)","notes":"EN/FR/DE/ES/IT business audio; not available for live."}],"transports":["WebSocket"],"audio":{"input":"Declare encoding (e.g. wav/pcm), sample_rate, bit_depth and channels at session init; they must match the chunks you send. Multi-channel supported.","output":"n/a"},"languages":"100+ with code-switching (Solaria-1).","latency":"No vendor latency number captured in this pass.","features":["partial transcripts","endpointing (silence threshold and max utterance duration)","code-switching","custom vocabulary","realtime translation, summarization, NER and other audio intelligence add-ons","multi-channel tagging","webhooks/callbacks","recorded audio retrievable after the session"],"pricing":{"model":"per-hour","items":[{"what":"Real-time, Starter (pay as you go)","price":"$0.75","unit":"per hour","notes":"Pricing page lists it as a 'starting at' price."},{"what":"Real-time, Growth","price":"as low as $0.25","unit":"per hour","notes":"Requires upfront commitment; exact rate not published."},{"what":"Enterprise","price":"custom","unit":"","notes":"Custom models, zero data retention option."}],"est_per_minute_usd":{"low":0.0042,"high":0.0125,"basis":"$0.25/hr Growth floor to $0.75/hr Starter."},"free_tier":"One-time EUR 50 credit (vendor estimates 60+ realtime hours); does not renew.","source":"https://www.gladia.io/pricing"},"limits":["Concurrent live sessions: Free 1, Paid 30 (default), Enterprise on demand; HTTP 429 when exhausted.","A single live session cannot exceed 3 hours.","Paid accounts can also queue up to 300 async jobs (not relevant to live)."],"regions":"Not re-verified in this pass.","setup":{"steps":["Get an API key (x-gladia-key).","POST https://api.gladia.io/v2/live with encoding, sample_rate, bit_depth, channels and options; the response returns a session id and a WebSocket url.","Connect to that url (the token is embedded) and send binary audio chunks.","Read transcript messages (is_final flags partial vs final), send {\"type\":\"stop_recording\"} to end, then GET /v2/live/{id} for the full result."],"endpoint":"POST https://api.gladia.io/v2/live, then the wss url returned in the response","auth":"x-gladia-key header on the init POST only; the returned WebSocket URL carries its own session token, so browsers can connect without the API key.","snippet_lang":"python","snippet":"import asyncio, json, os, requests, websockets  # pip install requests websockets>=14\n\ncfg = {\"encoding\": \"wav/pcm\", \"sample_rate\": 16000, \"bit_depth\": 16, \"channels\": 1,\n       \"messages_config\": {\"receive_partial_transcripts\": True}}\nr = requests.post(\"https://api.gladia.io/v2/live\", json=cfg,\n                  headers={\"x-gladia-key\": os.environ[\"GLADIA_API_KEY\"]})\nr.raise_for_status()\nws_url = r.json()[\"url\"]  # session-scoped, safe to hand to a browser\n\nasync def main():\n    async with websockets.connect(ws_url) as ws:\n        async def send():\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:\n                while chunk := f.read(3200):  # 100 ms\n                    await ws.send(chunk)\n                    await asyncio.sleep(0.1)\n            await ws.send(json.dumps({\"type\": \"stop_recording\"}))\n        asyncio.create_task(send())\n        async for msg in ws:\n            d = json.loads(msg)\n            if d.get(\"type\") == \"transcript\":\n                kind = \"FINAL\" if d[\"data\"][\"is_final\"] else \"partial\"\n                print(kind, d[\"data\"][\"utterance\"][\"text\"])\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Audio and transcripts are stored by default","detail":"Paid accounts keep audio input and transcripts for 3 weeks by default; free accounts keep everything for 1 year. Zero data retention is Enterprise-only and disables upload and polling endpoints."},{"severity":"medium","title":"Free tier allows only 1 live session","detail":"Free accounts get 1 concurrent live session (paid default 30), so any parallel testing needs a paid plan."},{"severity":"medium","title":"Diarization documented for pre-recorded","detail":"The speaker diarization page is for pre-recorded audio. Do not assume live speaker labels; use multi-channel input to separate parties on calls."},{"severity":"medium","title":"3-hour session cap","detail":"Live sessions are terminated after 3 hours. Long events need a new session and stitching of transcripts."},{"severity":"low","title":"Format declared up front","detail":"encoding, sample_rate, bit_depth and channels must exactly match what you send or the transcript degrades or fails; there is no auto-detection for raw PCM."}],"best_for":"Multilingual live transcription with code-switching and EU vendor preference, with optional realtime translation.","open_source":false,"self_hostable":false,"compliance":"Default retention: paid 3 weeks for audio/transcripts, 1 year metadata; Enterprise custom or zero retention. Certifications not re-verified.","docs":[{"label":"Pricing","url":"https://www.gladia.io/pricing"},{"label":"Models","url":"https://docs.gladia.io/chapters/introduction/models"},{"label":"Live quickstart","url":"https://docs.gladia.io/chapters/live-stt/quickstart"},{"label":"Concurrency and limits","url":"https://docs.gladia.io/chapters/limits-and-specifications/concurrency"},{"label":"Data retention","url":"https://docs.gladia.io/chapters/limits-and-specifications/data-retention"}],"sources":["https://www.gladia.io/pricing","https://docs.gladia.io/chapters/introduction/models.md","https://docs.gladia.io/chapters/limits-and-specifications/concurrency.md","https://docs.gladia.io/chapters/limits-and-specifications/data-retention.md","https://docs.gladia.io/chapters/live-stt/quickstart.md"],"confidence":"medium","unverified":"Exact WebSocket message field names in the snippet (data.is_final, data.utterance.text, stop_recording) are from the v2 API as previously documented and were not re-read line by line. Region and latency info not captured.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":100,"max_session_min":180,"concurrency":30,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.75,"diarization_streaming":null,"turn_detection":true,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":true,"telephony_8k":null,"bills_silence":null,"_notes":"100+ languages (Solaria-1). Free credit is EUR 50, one-time. $0.75/hr is the Starter PAYG price; Growth as low as $0.25/hr with commitment. Diarization documented for pre-recorded only. Audio and transcripts retained 3 weeks by default on paid plans."}},{"id":"rev-ai","name":"Rev AI Streaming","vendor":"Rev","category":"stt","summary":"Simple WebSocket streaming API with access token in the URL and custom vocabulary. Billing counts the longer of stream time and audio time.","status":"GA","models":[{"name":"Reverb (streaming)","status":"GA","notes":"Docs do not name the streaming model; the 'transcriber' parameter selects options. Treat model naming as unverified."}],"transports":["WebSocket"],"audio":{"input":"content_type query param, e.g. audio/x-raw;layout=interleaved;rate=16000;format=S16LE;channels=1. Raw, FLAC or WAV recommended.","output":"n/a"},"languages":"language param defaults to en; streaming language list not shown on the API page.","latency":"No vendor figure captured.","features":["partial and final hypotheses","custom vocabulary (custom_vocabulary_id)","profanity filter / options via query params"],"pricing":{"model":"per-hour","items":[{"what":"Reverb Transcription (English)","price":"$0.20","unit":"per hour","notes":"Pricing page does not say whether this line applies to streaming."},{"what":"Reverb Foreign Language","price":"$0.30","unit":"per hour","notes":"57 other languages on the pricing page (streaming coverage not stated)."},{"what":"Billing rule (streaming)","price":"max(stream duration, audio duration)","unit":"rounded up per second, 15 s minimum","notes":"From the streaming API docs."}],"est_per_minute_usd":{"low":0.0033,"high":0.005,"basis":"Assumes the $0.20-$0.30/hr Reverb lines apply to streaming, which the pricing page does not confirm."},"free_tier":"Free credits equal to 5 hours of Reverb ASR on Pay As You Go.","source":"https://www.rev.ai/pricing"},"limits":["3 hours per stream; open a new connection before the limit.","Default streaming concurrency 10 (support can raise).","max_connection_wait_seconds default 60."],"regions":"Rev markets 'global deployments' (EU option); not re-verified.","setup":{"steps":["Generate an access token in the Rev AI dashboard.","Open wss://api.rev.ai/speechtotext/v1/stream?access_token=...&content_type=...","Wait for the 'connected' message, then send binary audio.","Read 'partial' and 'final' messages; send the text 'EOS' to end."],"endpoint":"wss://api.rev.ai/speechtotext/v1/stream","auth":"access_token query parameter (do not expose the long-lived token in browsers; proxy through your server).","snippet_lang":"python","snippet":"import asyncio, json, os, websockets  # pip install websockets>=14\nfrom urllib.parse import urlencode\n\nq = urlencode({\"access_token\": os.environ[\"REVAI_ACCESS_TOKEN\"],\n               \"content_type\": \"audio/x-raw;layout=interleaved;rate=16000;format=S16LE;channels=1\"})\nURL = \"wss://api.rev.ai/speechtotext/v1/stream?\" + q\n\nasync def main():\n    async with websockets.connect(URL) as ws:\n        async def send():\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:\n                while chunk := f.read(3200):  # 100 ms\n                    await ws.send(chunk)\n                    await asyncio.sleep(0.1)\n            await ws.send(\"EOS\")\n        async for msg in ws:\n            d = json.loads(msg)\n            if d[\"type\"] == \"connected\":\n                asyncio.create_task(send())\n            elif d[\"type\"] == \"partial\":\n                print(\"partial\", \" \".join(e[\"value\"] for e in d[\"elements\"]))\n            elif d[\"type\"] == \"final\":\n                print(\"FINAL\", \"\".join(e[\"value\"] for e in d[\"elements\"]))\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Idle connection time is billed","detail":"You are charged for the larger of stream duration and audio duration, rounded up per second with a 15-second minimum per stream. An open socket with no audio still costs money."},{"severity":"medium","title":"Streaming price not itemised","detail":"rev.ai/pricing lists Reverb at $0.20/hr and foreign languages at $0.30/hr but never labels a streaming line. Confirm the streaming rate with Rev before budgeting."},{"severity":"medium","title":"Low default concurrency","detail":"Streaming concurrency defaults to 10 connections per account; request an increase before launch."},{"severity":"low","title":"Token in the URL","detail":"The access token goes in the query string, so it can leak into proxy and server logs. Use a server-side relay rather than putting it in browser code."}],"best_for":"Teams already on Rev for batch/human transcription who want a straightforward English streaming feed.","open_source":false,"self_hostable":false,"compliance":"Rev markets HIPAA support (site navigation lists HIPAA); not re-verified.","docs":[{"label":"Streaming API","url":"https://docs.rev.ai/api/streaming/"},{"label":"Pricing","url":"https://www.rev.ai/pricing"}],"sources":["https://docs.rev.ai/api/streaming/","https://www.rev.ai/pricing"],"confidence":"medium","unverified":"Streaming-specific price, streaming language list and current model name. Message field names in the snippet follow the documented partial/final element format but were not re-read in full.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":180,"concurrency":10,"free_tier":true,"free_credit_usd":null,"hipaa":true,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.2,"diarization_streaming":null,"turn_detection":null,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":null,"telephony_8k":null,"bills_silence":true,"_notes":"Free credit is 5 hours of Reverb ASR, not a dollar amount. $0.20/hr is the English Reverb line; pricing page does not confirm it applies to streaming. Billed on max(stream, audio) duration with a 15 s minimum. HIPAA is a vendor claim, not re-verified. Streaming language list not published."}},{"id":"soniox","name":"Soniox Real-time STT (stt-rt-v5)","vendor":"Soniox","category":"stt","summary":"Very cheap token-priced realtime STT over WebSocket with 60+ languages, speaker separation, language ID, endpoint detection and realtime translation included in one model.","status":"GA","models":[{"name":"stt-rt-v5","status":"Active (released 2026-06-16)","notes":"Current realtime model. Alias stt-rt-v4 now points here; v4 removed 2026-06-30 and auto-routed to v5."},{"name":"stt-async-v5","status":"Active","notes":"File/async counterpart, same accuracy and features per pricing page."}],"transports":["WebSocket"],"audio":{"input":"audio_format 'auto' for containers, or raw PCM: pcm_s8/s16/s24/s32, unsigned u8-u32, float f32/f64 (le/be), with sample_rate and num_channels.","output":"n/a"},"languages":"60+ languages with language identification and mixed-language robustness; realtime translation across 3,600+ language pairs (vendor).","latency":"No numeric vendor claim captured; model notes cite 'faster semantic endpointing'.","features":["non-final and final tokens","semantic endpoint detection (<end> token, tunable sensitivity and max delay)","manual finalize (<fin> token)","speaker separation (diarization) included","language identification included","context / custom vocabulary","realtime translation","keepalive control message","temporary API keys for browsers"],"pricing":{"model":"per-token","items":[{"what":"Real-time input audio","price":"$2.00","unit":"per 1M audio tokens","notes":"About 30,000 audio tokens per hour of audio."},{"what":"Real-time input text (context)","price":"$4.00","unit":"per 1M tokens","notes":""},{"what":"Real-time output text","price":"$4.00","unit":"per 1M tokens","notes":"About 15,000 output tokens per hour of speech."},{"what":"Real-time effective rate","price":"~$0.12","unit":"per hour","notes":"Vendor calculator; diarization, language ID and formatting included. Async is ~$0.10/hr."}],"est_per_minute_usd":{"low":0.002,"high":0.0025,"basis":"~$0.12/hr vendor estimate = $0.002/min; higher with long context prompts or translation output."},"free_tier":"Not stated on the pricing page.","source":"https://soniox.com/pricing"},"limits":["100 requests per minute, 10 concurrent streams by default (raisable in console).","Each stream capped at 300 minutes; this cap cannot be raised.","Open-connection limit is a multiple of the concurrent-request limit, including idle sockets."],"regions":"Not re-verified in this pass.","setup":{"steps":["Create an API key in the Soniox Console.","Connect to wss://stt-rt.soniox.com/transcribe-websocket with Authorization: Bearer <key> (or a temporary key for browsers).","Send a JSON start message with model, audio_format, sample_rate, num_channels and options.","Stream binary audio; rebuild text from tokens (append is_final tokens, redraw non-final ones).","Send an empty TEXT frame to finish; wait for {\"finished\": true}."],"endpoint":"wss://stt-rt.soniox.com/transcribe-websocket","auth":"Authorization: Bearer <API_KEY> header (api_key in the start message is deprecated); temporary API keys for client-side use.","snippet_lang":"python","snippet":"import asyncio, json, os, websockets  # pip install websockets>=14\n\nURL = \"wss://stt-rt.soniox.com/transcribe-websocket\"\nCFG = {\"model\": \"stt-rt-v5\", \"audio_format\": \"pcm_s16le\", \"sample_rate\": 16000,\n       \"num_channels\": 1, \"enable_endpoint_detection\": True}\n\nasync def main():\n    hdr = {\"Authorization\": f\"Bearer {os.environ['SONIOX_API_KEY']}\"}\n    async with websockets.connect(URL, additional_headers=hdr) as ws:\n        await ws.send(json.dumps(CFG))\n        async def send():\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:\n                while chunk := f.read(3200):  # 100 ms\n                    await ws.send(chunk)\n                    await asyncio.sleep(0.1)\n            await ws.send(\"\")  # empty TEXT frame ends the stream\n        asyncio.create_task(send())\n        final = \"\"\n        async for msg in ws:\n            d = json.loads(msg)\n            if d.get(\"error_code\"):\n                print(d); break\n            toks = d.get(\"tokens\", [])\n            final += \"\".join(t[\"text\"] for t in toks if t[\"is_final\"])\n            partial = \"\".join(t[\"text\"] for t in toks if not t[\"is_final\"])\n            print(final + \" | \" + partial)\n            if d.get(\"finished\"):\n                break\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Default concurrency is only 10","detail":"New accounts get 10 concurrent streams and 100 requests per minute. Raise it in the console well before a launch."},{"severity":"medium","title":"Hard 300-minute stream cap","detail":"Each session ends at 300 minutes and Soniox says this cannot be increased; long broadcasts must roll over to a new session."},{"severity":"medium","title":"Empty binary frame does not end the stream","detail":"Only an empty TEXT frame finishes the session; an empty binary frame is treated as an empty audio chunk and the session stays open."},{"severity":"medium","title":"Token-based bill is an estimate","detail":"Pricing is per audio, context and output token; the ~$0.12/hr figure assumes typical speech density. Large context prompts and translation output add text tokens."},{"severity":"low","title":"Model auto-upgrades via aliases","detail":"stt-rt-v4 was silently routed to stt-rt-v5 after 2026-06-30. If you need stable behaviour for evaluation, pin the exact model and watch the changelog."}],"best_for":"Cost-sensitive multilingual live transcription and live translation, with diarization and language ID without add-on fees.","open_source":false,"self_hostable":false,"compliance":"Not re-verified in this pass.","docs":[{"label":"Pricing","url":"https://soniox.com/pricing"},{"label":"Models and changelog","url":"https://soniox.com/docs/stt/models"},{"label":"Real-time transcription","url":"https://soniox.com/docs/stt/rt/real-time-transcription"},{"label":"RT limits","url":"https://soniox.com/docs/stt/rt/limits-and-quotas"},{"label":"WebSocket API reference","url":"https://soniox.com/docs/api-reference/stt/websocket-api"}],"sources":["https://soniox.com/pricing","https://soniox.com/docs/stt/models.mdx","https://soniox.com/docs/stt/rt/limits-and-quotas.mdx","https://soniox.com/docs/stt/rt/real-time-transcription.mdx","https://soniox.com/docs/api-reference/stt/websocket-api.mdx"],"confidence":"high","unverified":"Free credit, regions and compliance.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"Soniox","facts":{"cat":"stt","latency_ms":null,"languages":60,"max_session_min":300,"concurrency":10,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.12,"diarization_streaming":true,"turn_detection":true,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":true,"telephony_8k":null,"bills_silence":null,"_notes":"Token-billed; $0.12/hr is the vendor calculator estimate with diarization and language ID included. 60+ languages. 300-minute stream cap cannot be raised. Concurrency raisable in console."}},{"id":"openai-realtime-transcription","name":"OpenAI Realtime transcription sessions","vendor":"OpenAI","category":"stt","summary":"Transcription-only sessions on the Realtime API (WebSocket or WebRTC). gpt-live-transcribe and gpt-realtime-whisper stream deltas; gpt-transcribe and gpt-4o-transcribe transcribe each committed turn.","status":"GA","models":[{"name":"gpt-live-transcribe","status":"GA (recommended in realtime guide)","notes":"Streams deltas, tunable delay (minimal..xhigh), prompt, keywords, multiple language hints. Does NOT support server_vad/semantic_vad: you must commit turns yourself. No language detection output."},{"name":"gpt-realtime-whisper","status":"GA","notes":"Streaming STT priced by audio duration; supported on v1/realtime/transcription_sessions."},{"name":"gpt-transcribe","status":"GA","notes":"High-accuracy model for committed turns (starts after commit) and files; returns detected languages."},{"name":"gpt-4o-transcribe","status":"GA (older)","notes":"Token-priced; usable as transcription model in realtime sessions."},{"name":"gpt-4o-mini-transcribe","status":"GA (older)","notes":"Cheapest token-priced option."},{"name":"gpt-4o-transcribe-diarize","status":"GA (file transcription)","notes":"Speaker labels; file API, not a live streaming model."}],"transports":["WebSocket","WebRTC"],"audio":{"input":"audio/pcm at 24 kHz in the documented example (16-bit, base64 in input_audio_buffer.append); G.711 formats historically supported on Realtime (not re-verified).","output":"n/a"},"languages":"Multilingual (count not published on the pages read); gpt-live-transcribe accepts a 'languages' hint list.","latency":"Vendor: 'low-latency' with tunable delay setting; no number published.","features":["transcript deltas","final transcript per committed turn","client-side commit (input_audio_buffer.commit)","prompt / keywords / language hints","delay setting (minimal to xhigh) on gpt-live-transcribe","WebRTC for browsers with ephemeral client secrets"],"pricing":{"model":"per-minute","items":[{"what":"gpt-live-transcribe","price":"$0.017","unit":"per minute of audio","notes":"Realtime audio duration pricing."},{"what":"gpt-realtime-whisper","price":"$0.017","unit":"per minute of audio","notes":""},{"what":"gpt-transcribe","price":"$0.0045","unit":"per minute","notes":"Per committed turn / file."},{"what":"gpt-4o-transcribe","price":"$2.50 in / $10.00 out","unit":"per 1M tokens","notes":"OpenAI estimate ~$0.006/min."},{"what":"gpt-4o-mini-transcribe","price":"$1.25 in / $5.00 out","unit":"per 1M tokens","notes":"OpenAI estimate ~$0.003/min."},{"what":"Whisper (file API)","price":"$0.006","unit":"per minute","notes":"Batch only, not streaming."}],"est_per_minute_usd":{"low":0.003,"high":0.017,"basis":"$0.003/min (4o-mini-transcribe estimate, per turn) to $0.017/min (true streaming deltas with gpt-live-transcribe or gpt-realtime-whisper)."},"free_tier":"None specific to transcription.","source":"https://developers.openai.com/api/docs/pricing"},"limits":["Rate limits depend on account usage tier (not captured).","Completion events from different turns can arrive out of order; match by item_id."],"regions":"OpenAI API global; data residency options not checked for these models.","setup":{"steps":["Create an API key.","Server: open a Realtime WebSocket with Authorization: Bearer <key>; the OpenAI cookbook pattern uses wss://api.openai.com/v1/realtime?intent=transcription.","Send session.update with session.type='transcription', audio.input.format, transcription.model and turn_detection (null for gpt-live-transcribe).","Append base64 PCM with input_audio_buffer.append, commit each turn with input_audio_buffer.commit (your own VAD).","Browser: mint an ephemeral client secret on your server and connect via WebRTC."],"endpoint":"wss://api.openai.com/v1/realtime?intent=transcription (route listed as v1/realtime/transcription_sessions on model pages)","auth":"Authorization: Bearer <API_KEY> server side; ephemeral client secrets for browser WebRTC.","snippet_lang":"javascript","snippet":"import WebSocket from \"ws\"; // npm i ws\nimport fs from \"fs\";\n\nconst ws = new WebSocket(\"wss://api.openai.com/v1/realtime?intent=transcription\", {\n  headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },\n});\nws.on(\"open\", async () => {\n  ws.send(JSON.stringify({ type: \"session.update\", session: { type: \"transcription\",\n    audio: { input: { format: { type: \"audio/pcm\", rate: 24000 },\n      transcription: { model: \"gpt-live-transcribe\" }, turn_detection: null } } } }));\n  const pcm = fs.readFileSync(\"audio_24k_mono_s16le.raw\");\n  for (let i = 0; i < pcm.length; i += 4800) { // 100 ms at 24 kHz\n    ws.send(JSON.stringify({ type: \"input_audio_buffer.append\",\n      audio: pcm.subarray(i, i + 4800).toString(\"base64\") }));\n    await new Promise((r) => setTimeout(r, 100));\n  }\n  ws.send(JSON.stringify({ type: \"input_audio_buffer.commit\" })); // end of turn\n});\nws.on(\"message\", (data) => {\n  const ev = JSON.parse(data);\n  if (ev.type === \"conversation.item.input_audio_transcription.delta\") process.stdout.write(ev.delta);\n  if (ev.type === \"conversation.item.input_audio_transcription.completed\") console.log(\"\\nFINAL:\", ev.transcript);\n  if (ev.type === \"error\") console.error(ev.error);\n});"},"warnings":[{"severity":"high","title":"No server VAD on gpt-live-transcribe","detail":"The recommended streaming model rejects server_vad and semantic_vad. You must run client-side VAD and send input_audio_buffer.commit at the end of each turn, or you never get a final transcript."},{"severity":"high","title":"True streaming is ~3x the cost of competitors","detail":"Delta-streaming models (gpt-live-transcribe, gpt-realtime-whisper) are $0.017/min, roughly 3-8x Deepgram, AssemblyAI or Soniox list prices. gpt-transcribe at $0.0045/min only returns text after a commit."},{"severity":"medium","title":"Model names churned in 2026","detail":"gpt-realtime-whisper (May 2026) and gpt-live-transcribe / gpt-transcribe (July 2026) arrived within months; third-party guides disagree on which to use. Follow the realtime transcription guide and pin a model."},{"severity":"medium","title":"Out-of-order completions","detail":"Completed events from different turns can arrive out of order. Key transcripts by item_id rather than appending in arrival order."},{"severity":"medium","title":"Azure lags OpenAI","detail":"A May 2026 Microsoft Q&A thread reports Azure did not list gpt-realtime-whisper for live input transcription while it worked on OpenAI directly. Check model availability per Azure region."},{"severity":"low","title":"Language param is plural","detail":"gpt-live-transcribe takes 'languages' (a list); sending the older singular 'language' field is not supported."}],"best_for":"Apps already on the OpenAI Realtime stack that want transcription in the same session model, or browser capture via WebRTC.","open_source":false,"self_hostable":false,"compliance":"OpenAI API data controls apply (API data not used for training by default per OpenAI policy); BAA availability not re-verified for these models.","docs":[{"label":"Realtime transcription guide","url":"https://developers.openai.com/api/docs/guides/realtime-transcription"},{"label":"Pricing","url":"https://developers.openai.com/api/docs/pricing"},{"label":"gpt-live-transcribe model","url":"https://developers.openai.com/api/docs/models/gpt-live-transcribe"},{"label":"gpt-realtime-whisper model","url":"https://developers.openai.com/api/docs/models/gpt-realtime-whisper"},{"label":"gpt-transcribe model","url":"https://developers.openai.com/api/docs/models/gpt-transcribe"}],"sources":["https://developers.openai.com/api/docs/pricing","https://developers.openai.com/api/docs/guides/realtime-transcription.md","https://developers.openai.com/api/docs/models/gpt-live-transcribe.md","https://developers.openai.com/api/docs/models/gpt-realtime-whisper.md","https://developers.openai.com/api/docs/models/gpt-transcribe.md","https://learn.microsoft.com/en-us/answers/a/12769224","https://www.datacamp.com/tutorial/gpt-live-transcribe-api"],"confidence":"medium","unverified":"Exact WebSocket URL for transcription sessions (the ?intent=transcription query comes from a cookbook pattern, not the guide), rate limits, supported language count, and whether gpt-realtime-whisper is being superseded by gpt-live-transcribe.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"OpenAI transcription","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":1.02,"diarization_streaming":false,"turn_detection":null,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":null,"telephony_8k":null,"bills_silence":null,"_notes":"Price is true delta streaming (gpt-live-transcribe or gpt-realtime-whisper, $0.017/min). gpt-transcribe is $0.27/hr but only returns text after a commit. gpt-live-transcribe has no server VAD; you commit turns yourself. Diarization model is file-only. Rate limits depend on account tier."}},{"id":"google-cloud-stt-v2","name":"Google Cloud Speech-to-Text V2 streaming (Chirp 3)","vendor":"Google Cloud","category":"stt","summary":"gRPC StreamingRecognize on Speech-to-Text V2 with the chirp_3 model in us/eu multi-regions. Strong language coverage, but short stream limit and no streaming diarization on Chirp 3.","status":"GA","models":[{"name":"chirp_3","status":"GA (us and eu multi-regions)","notes":"StreamingRecognize supported. Diarization only in BatchRecognize. Speech adaptation up to 1,000 phrases. Language-agnostic auto-detect GA."},{"name":"chirp_2","status":"Older generation","notes":"Not re-checked in this pass."},{"name":"telephony / long / short / latest_*","status":"Legacy standard models","notes":"Priced as 'Standard' in the V2 table."}],"transports":["gRPC"],"audio":{"input":"ExplicitDecodingConfig (LINEAR16, MULAW, ALAW and others) with sample rate and channel count, or auto-decoding for containers. Audio must be sent at roughly real-time pace.","output":"n/a"},"languages":"Chirp 3 table lists about 111 locales (my count: 29 GA, 82 Preview).","latency":"No vendor figure captured.","features":["streaming interim results (check per model)","utterance-level timestamps (streaming only)","speech adaptation / phrase sets","language-agnostic auto-detect","voice activity events"],"pricing":{"model":"per-minute","items":[{"what":"V2 Standard recognition (incl. chirp)","price":"$0.016","unit":"per minute, 0-500K min/month","notes":"Same SKU for streaming and sync. 1-second rounding per request."},{"what":"V2 Standard, 500K-1M","price":"$0.010","unit":"per minute","notes":""},{"what":"V2 Standard, 1M-2M","price":"$0.008","unit":"per minute","notes":""},{"what":"V2 Standard, 2M+","price":"$0.004","unit":"per minute","notes":""},{"what":"V2 Dynamic batch (not streaming)","price":"$0.003","unit":"per minute","notes":"Discounted low-urgency batch."},{"what":"V1 API","price":"$0.016 with data logging / $0.024 without","unit":"per minute after 60 free min/month","notes":"V1 only. Medical models $0.078/min."}],"est_per_minute_usd":{"low":0.004,"high":0.016,"basis":"V2 tiered price; most teams pay $0.016/min until 500K minutes per month. Multi-channel audio is billed per channel."},"free_tier":"V1 includes 60 free minutes/month; new Google Cloud accounts get $300 trial credit (Google Cloud free program).","source":"https://cloud.google.com/speech-to-text/pricing"},"limits":["Streaming requests limited to about 5 minutes of audio per stream; reconnect and stitch for longer audio.","10 MB limit applies to StreamingRecognize request messages.","Phrase sets: 5,000 phrases / 100,000 characters per request (general quotas page); Chirp 3 adaptation dictionary up to 1,000 phrases."],"regions":"chirp_3: us and eu multi-regions (GA); more planned. Use the regional API endpoint (e.g. us-speech.googleapis.com) matching the recognizer location.","setup":{"steps":["Enable the Speech-to-Text API on a Google Cloud project and authenticate with a service account (ADC).","pip install google-cloud-speech.","Create a SpeechClient with api_endpoint set to the regional host (us-speech.googleapis.com for location us).","Send a first StreamingRecognizeRequest with recognizer projects/<id>/locations/us/recognizers/_ and streaming_config, then audio-only requests.","Restart the stream before ~5 minutes."],"endpoint":"gRPC us-speech.googleapis.com:443 / eu-speech.googleapis.com:443 (Speech.StreamingRecognize, V2)","auth":"Google Cloud IAM (service account / ADC OAuth tokens). No browser-direct option; proxy through your backend.","snippet_lang":"python","snippet":"from google.api_core.client_options import ClientOptions\nfrom google.cloud.speech_v2 import SpeechClient\nfrom google.cloud.speech_v2.types import cloud_speech as cs  # pip install google-cloud-speech\n\nPROJECT = \"my-project\"\nclient = SpeechClient(client_options=ClientOptions(api_endpoint=\"us-speech.googleapis.com\"))\ncfg = cs.RecognitionConfig(\n    explicit_decoding_config=cs.ExplicitDecodingConfig(\n        encoding=cs.ExplicitDecodingConfig.AudioEncoding.LINEAR16,\n        sample_rate_hertz=16000, audio_channel_count=1),\n    language_codes=[\"en-US\"], model=\"chirp_3\")\nscfg = cs.StreamingRecognitionConfig(\n    config=cfg, streaming_features=cs.StreamingRecognitionFeatures(interim_results=True))\n\ndef requests():\n    yield cs.StreamingRecognizeRequest(\n        recognizer=f\"projects/{PROJECT}/locations/us/recognizers/_\", streaming_config=scfg)\n    with open(\"audio_16k_mono.raw\", \"rb\") as f:\n        while chunk := f.read(3200):  # 100 ms; send at real-time pace for live audio\n            yield cs.StreamingRecognizeRequest(audio=chunk)\n\nfor resp in client.streaming_recognize(requests=requests()):\n    for r in resp.results:\n        if r.alternatives:\n            print(\"FINAL\" if r.is_final else \"partial\", r.alternatives[0].transcript)"},"warnings":[{"severity":"high","title":"About 5 minutes per stream","detail":"Streaming requests are limited to roughly 5 minutes of audio. Long calls need you to open a new stream before the limit and stitch results, handling words cut at the boundary."},{"severity":"high","title":"No streaming diarization on Chirp 3","detail":"Chirp 3 speaker diarization is only available in BatchRecognize. Word-level timestamps are also not supported in streaming."},{"severity":"medium","title":"Region-locked endpoint","detail":"chirp_3 is only in the us and eu multi-regions; the recognizer location and the API endpoint host must match or calls fail."},{"severity":"medium","title":"Expensive at low volume","detail":"$0.016/min until 500K minutes a month is 3x Deepgram Nova-3 PAYG. Discounts only kick in at large volumes."},{"severity":"low","title":"Many Chirp 3 locales are Preview","detail":"Of the ~111 listed locales most are Preview, which carries no SLA; check your language's status before committing."}],"best_for":"GCP-native apps that need wide language coverage and IAM/VPC-SC controls, with short utterances (commands, IVR).","open_source":false,"self_hostable":false,"compliance":"Google Cloud compliance programs (HIPAA BAA via Google Cloud, etc.) apply per Google's covered-services list; data logging opt-in affects V1 pricing. Not re-verified per model.","docs":[{"label":"Pricing","url":"https://cloud.google.com/speech-to-text/pricing"},{"label":"Chirp 3 model","url":"https://docs.cloud.google.com/speech-to-text/v2/docs/chirp_3-model"},{"label":"Quotas and limits","url":"https://docs.cloud.google.com/speech-to-text/quotas"}],"sources":["https://cloud.google.com/speech-to-text/pricing","https://docs.cloud.google.com/speech-to-text/v2/docs/chirp_3-model","https://docs.cloud.google.com/speech-to-text/quotas"],"confidence":"medium","unverified":"Whether chirp_3 returns interim results in streaming, chirp_2 status, and any V2 free minutes. The V2 pricing table was read from page markup.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"Google Cloud STT","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":5,"concurrency":null,"free_tier":true,"free_credit_usd":300,"hipaa":true,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":false,"grpc":true,"self_hostable":false,"open_weights":false,"price_per_hour":0.96,"diarization_streaming":false,"turn_detection":null,"keyterms":true,"pii_redaction":null,"interim_results":null,"code_switching":null,"telephony_8k":true,"bills_silence":null,"_notes":"Chirp 3 lists about 111 locales (29 GA, 82 Preview), not a language count. $0.016/min V2 standard up to 500K min/month, cheaper at volume. About 5 minutes of audio per stream. $300 is the general Google Cloud trial credit. HIPAA via Google Cloud BAA, not re-verified per model. Chirp 3 diarization is batch only; interim results not confirmed for chirp_3."}},{"id":"azure-ai-speech-realtime","name":"Azure AI Speech real-time speech to text","vendor":"Microsoft","category":"stt","summary":"Mature real-time STT via the Speech SDK (WebSocket under the hood) with intermediate results, phrase lists, custom speech models, continuous language ID and real-time diarization add-ons.","status":"GA","models":[{"name":"Standard (base) real-time model","status":"GA","notes":"Latest base model per locale used by default."},{"name":"Custom speech endpoint","status":"GA","notes":"Trained/adapted model, needs hosted endpoint (hourly hosting fee)."},{"name":"MAI-Transcribe-2 / MAI-Transcribe-1.5","status":"Listed on pricing page","notes":"Priced per hour; MAI-Transcribe-2 under a limited-time promotion to 2026-12-31. Real-time availability not confirmed."},{"name":"Multichannel real-time transcription","status":"Preview","notes":"Up to two channels tagged independently."}],"transports":["WebSocket (via Speech SDK)","REST for short audio"],"audio":{"input":"Speech SDK handles microphone, files and push/pull streams; default 16 kHz 16-bit mono PCM; compressed formats via GStreamer.","output":"n/a"},"languages":"100+ locales (see language support page; count not re-verified).","latency":"No vendor figure captured.","features":["intermediate (recognizing) results","continuous recognition","phrase lists","custom speech models","continuous language identification (paid add-on)","real-time diarization (paid add-on)","pronunciation assessment","multichannel (preview)"],"pricing":{"model":"per-hour","items":[{"what":"Standard real-time STT (S1, eastus)","price":"$1.00","unit":"per audio hour","notes":"Azure retail prices API meter 'S1 Speech To Text'."},{"what":"Custom real-time STT","price":"$1.20","unit":"per audio hour","notes":"Plus custom endpoint hosting ~$0.0538/hr per model."},{"what":"Real-time enhanced features add-on (diarization / language ID)","price":"$0.30","unit":"per audio hour per feature","notes":"Meter 'S1 Speech to Text Enhanced Feature Audio'."},{"what":"Commitment tiers","price":"2K / 10K / 50K / 100K hour bundles","unit":"monthly","notes":"Overage $0.80/hr (2K), $0.65 (10K), $0.50 (50K), $0.40 (100K) for standard STT."},{"what":"Batch STT (reference)","price":"$0.18","unit":"per audio hour","notes":"Fast transcription $0.36/hr."}],"est_per_minute_usd":{"low":0.0067,"high":0.025,"basis":"$0.40/hr commitment-tier overage up to $1.20/hr custom + $0.30/hr add-on. Pay-as-you-go standard is $0.0167/min."},"free_tier":"F0: 5 audio hours per month shared between standard and custom real-time; 1 concurrent request.","source":"https://azure.microsoft.com/en-us/pricing/details/speech/"},"limits":["Concurrent real-time requests: F0 1 (fixed), S0 100 default (adjustable); limit is shared with speech translation.","Diarization identifies up to 35 speakers (errors beyond that)."],"regions":"Many Azure regions; prices above are eastus retail list prices.","setup":{"steps":["Create a Speech (or Foundry) resource and copy key + region.","pip install azure-cognitiveservices-speech.","Create SpeechConfig and a SpeechRecognizer; subscribe to recognizing (partial) and recognized (final) events.","Call start_continuous_recognition; for browsers fetch a short-lived authorization token from your backend (issueToken endpoint)."],"endpoint":"Speech SDK manages the WebSocket (wss://<region>.stt.speech.microsoft.com/...); use SDK rather than raw protocol.","auth":"Ocp-Apim-Subscription-Key / resource key, Microsoft Entra ID, or 10-minute authorization tokens for client apps.","snippet_lang":"python","snippet":"import os, time\nimport azure.cognitiveservices.speech as speechsdk  # pip install azure-cognitiveservices-speech\n\ncfg = speechsdk.SpeechConfig(subscription=os.environ[\"SPEECH_KEY\"],\n                             region=os.environ[\"SPEECH_REGION\"])\ncfg.speech_recognition_language = \"en-US\"\naudio = speechsdk.audio.AudioConfig(use_default_microphone=True)\nrec = speechsdk.SpeechRecognizer(speech_config=cfg, audio_config=audio)\n\nrec.recognizing.connect(lambda e: print(\"partial\", e.result.text))\nrec.recognized.connect(lambda e: print(\"FINAL\", e.result.text))\nrec.canceled.connect(lambda e: print(\"canceled\", e.cancellation_details))\n\nrec.start_continuous_recognition()\ntime.sleep(30)\nrec.stop_continuous_recognition()"},"warnings":[{"severity":"high","title":"Pricing page shows no numbers","detail":"The public Speech pricing page renders rates as placeholders unless a region/currency loads; the figures here come from the Azure retail prices API (eastus, USD). Re-check for your region."},{"severity":"medium","title":"Most expensive mainstream list price","detail":"$1.00 per audio hour pay-as-you-go is several times Deepgram, AssemblyAI or Soniox. Commitment tiers are needed to get near $0.40-0.50/hr."},{"severity":"medium","title":"Diarization and language ID cost extra","detail":"Real-time diarization and continuous language identification are billed as add-ons (about $0.30 per audio hour per feature)."},{"severity":"medium","title":"Free tier is 1 concurrent stream","detail":"F0 allows a single concurrent real-time request and the limit cannot be raised; S0 starts at 100."},{"severity":"low","title":"Custom models carry hosting fees","detail":"Custom speech endpoints bill hosting per model per hour even when idle; unused free-tier models are decommissioned after 7 days."}],"best_for":"Enterprises on Azure needing custom acoustic/language models, on-prem containers, and Microsoft compliance coverage.","open_source":false,"self_hostable":true,"compliance":"Azure compliance scope (HIPAA BAA, SOC, ISO) applies to Azure AI Speech per Microsoft; connected/disconnected containers available. Not re-verified.","docs":[{"label":"Speech to text overview","url":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-to-text"},{"label":"Quotas and limits","url":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-services-quotas-and-limits"},{"label":"Pricing","url":"https://azure.microsoft.com/en-us/pricing/details/speech/"},{"label":"Retail prices API","url":"https://prices.azure.com/api/retail/prices"}],"sources":["https://azure.microsoft.com/en-us/pricing/details/speech/","https://prices.azure.com/api/retail/prices?$filter=armRegionName eq 'eastus' and contains(productName,'Speech')","https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-to-text","https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-services-quotas-and-limits"],"confidence":"medium","unverified":"Mapping of the 'Enhanced Feature Audio' meter to diarization/language-ID add-ons is inferred from meter names; MAI-Transcribe real-time support and prices; locale count.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"Azure Speech STT","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":100,"free_tier":true,"free_credit_usd":null,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_hour":1,"diarization_streaming":true,"turn_detection":null,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":true,"telephony_8k":null,"bills_silence":null,"_notes":"100+ locales, count not verified. $1.00/hr is eastus PAYG from the retail price API; commitment tiers go down to $0.40/hr. Diarization and continuous language ID are +$0.30/hr add-ons each. Free tier F0 is 5 audio hours/month with 1 concurrent stream. Compliance per Azure scope, not re-verified. WebSocket via Speech SDK."}},{"id":"aws-transcribe-streaming","name":"Amazon Transcribe Streaming (incl. Medical and Call Analytics)","vendor":"AWS","category":"stt","summary":"Streaming STT over HTTP/2 or WebSocket with SigV4 auth, plus Transcribe Medical streaming and real-time Call Analytics. Current us-east-1 price list shows a flat $0.01/min for standard streaming.","status":"GA","models":[{"name":"Standard streaming","status":"GA","notes":"Partial results stabilization, custom vocabulary, custom language models, PII identification/redaction, speaker labels, channel identification."},{"name":"Transcribe Medical streaming","status":"GA","notes":"US English medical specialties; $0.075/min."},{"name":"Call Analytics streaming","status":"GA","notes":"Real-time call insights; tiered from $0.03/min."}],"transports":["HTTP/2","WebSocket"],"audio":{"input":"PCM signed 16-bit little-endian (not WAV), FLAC, or Opus in Ogg. 16 kHz recommended; 8 kHz telephony accepted. Chunks of 50-200 ms recommended; send zero-byte silence rather than pausing.","output":"n/a"},"languages":"Many streaming languages (count not re-verified); Medical is US English.","latency":"No vendor figure captured; latency depends on chunk size per docs.","features":["partial results with stabilization","custom vocabulary and vocabulary filters","custom language models (CLM, priced separately)","PII identification/redaction (paid)","speaker partitioning","channel identification","automatic language identification","toxicity detection (batch)","Call Analytics real-time insights"],"pricing":{"model":"per-minute","items":[{"what":"Standard streaming (us-east-1)","price":"$0.01","unit":"per minute ($0.0001667/s)","notes":"AWS price list SKU 'StreamingAudio'; billed per second. Older blog posts quote $0.024/min tiered."},{"what":"PII redaction streaming","price":"+$0.0024 / $0.0015 / $0.00102 / $0.00078","unit":"per minute by tier (0-250K, 250K-1M, 1M-5M, 5M+)","notes":"Free tier does not cover this."},{"what":"Custom language model streaming","price":"$0.006 first 250K","unit":"per minute ($0.0001/s), tiered down","notes":""},{"what":"Medical streaming","price":"$0.075","unit":"per minute ($0.00125/s)","notes":""},{"what":"Call Analytics streaming","price":"$0.030 / $0.0186 / $0.0138 / $0.0114","unit":"per minute by tier","notes":""},{"what":"Batch (reference)","price":"$0.006","unit":"per minute","notes":"Price list 'TranscribeAudio' $0.0001/s."}],"est_per_minute_usd":{"low":0.01,"high":0.075,"basis":"$0.01/min standard streaming up to $0.075/min Medical. Add PII redaction or CLM as needed."},"free_tier":"60 minutes per month for 12 months from first request (excludes PII redaction).","source":"https://aws.amazon.com/transcribe/pricing/"},"limits":["25 concurrent standard streams per region by default (adjustable); Medical and Call Analytics streams also 25.","Streams must be close to real time; LimitExceededException on overuse.","Billed in 1-second increments, no per-request minimum for transcription."],"regions":"Most commercial AWS regions plus GovCloud; FIPS endpoints in US regions.","setup":{"steps":["Create IAM credentials with transcribe:StartStreamTranscription.","Use an AWS SDK with event-stream support (JS v3 @aws-sdk/client-transcribe-streaming, Java, Go, or the Python amazon-transcribe package).","Start a stream with language, sample rate and encoding, push audio events, read TranscriptEvents (IsPartial flag).","Browsers: use WebSocket with a SigV4 presigned URL generated server-side, or Cognito credentials."],"endpoint":"transcribestreaming.<region>.amazonaws.com (HTTP/2 :443; WebSocket :8443 /stream-transcription-websocket with presigned URL)","auth":"AWS SigV4 (IAM). WebSocket uses a presigned URL (max 5 minute validity).","snippet_lang":"python","snippet":"import asyncio  # pip install amazon-transcribe\nfrom amazon_transcribe.client import TranscribeStreamingClient\nfrom amazon_transcribe.handlers import TranscriptResultStreamHandler\n\nclass Printer(TranscriptResultStreamHandler):\n    async def handle_transcript_event(self, event):\n        for r in event.transcript.results:\n            for alt in r.alternatives:\n                print(\"partial\" if r.is_partial else \"FINAL\", alt.transcript)\n\nasync def main():\n    client = TranscribeStreamingClient(region=\"us-east-1\")\n    stream = await client.start_stream_transcription(\n        language_code=\"en-US\", media_sample_rate_hz=16000, media_encoding=\"pcm\")\n    async def write():\n        with open(\"audio_16k_mono.raw\", \"rb\") as f:\n            while chunk := f.read(3200):  # 100 ms\n                await stream.input_stream.send_audio_event(audio_chunk=chunk)\n                await asyncio.sleep(0.1)\n        await stream.input_stream.end_stream()\n    await asyncio.gather(write(), Printer(stream.output_stream).handle_events())\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Old price figures circulate","detail":"Many articles quote $0.024/min tiered streaming. The current AWS price list (us-east-1) shows $0.0001667/s = $0.01/min for StreamingAudio. Prices vary by region; check the price list API for yours."},{"severity":"medium","title":"Send silence, not gaps","detail":"AWS says to send zero-byte PCM silence when there is no speech and to keep 50-200 ms chunks. Gaps or bursty sends raise latency or trigger errors."},{"severity":"medium","title":"Medical is 7.5x standard","detail":"Transcribe Medical streaming is $0.075/min and US English only; Call Analytics streaming starts at $0.03/min."},{"severity":"medium","title":"WAV headers are not PCM","detail":"Streaming PCM must be raw signed 16-bit little-endian. Sending a WAV file including its header corrupts the start of the transcript."},{"severity":"low","title":"Python SDK is a separate package","detail":"boto3 does not do streaming transcription; use the amazon-transcribe package or another SDK with HTTP/2 event streams."}],"best_for":"AWS-centric call centers and healthcare (Medical, Call Analytics) needing IAM, VPC and HIPAA-eligible services.","open_source":false,"self_hostable":false,"compliance":"Amazon Transcribe and Transcribe Medical are HIPAA-eligible under the AWS BAA per AWS (not re-verified here). AI services opt-out policy controls data use for training.","docs":[{"label":"Pricing","url":"https://aws.amazon.com/transcribe/pricing/"},{"label":"Streaming guide","url":"https://docs.aws.amazon.com/transcribe/latest/dg/streaming.html"},{"label":"Endpoints and quotas","url":"https://docs.aws.amazon.com/general/latest/gr/transcribe.html"},{"label":"Price list (us-east-1)","url":"https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/transcribe/current/us-east-1/index.json"}],"sources":["https://aws.amazon.com/transcribe/pricing/","https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/transcribe/current/us-east-1/index.json","https://docs.aws.amazon.com/transcribe/latest/dg/streaming.html","https://docs.aws.amazon.com/general/latest/gr/transcribe.html"],"confidence":"medium","unverified":"The flat $0.01/min streaming price is from the price list JSON and the page's worked example; the rendered tier table was not visible. PII tier 4 value inferred from the price list (5M+ = $0.000013/s). Streaming language count.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"Amazon Transcribe","facts":{"cat":"stt","latency_ms":null,"languages":77,"max_session_min":null,"concurrency":25,"free_tier":true,"free_credit_usd":null,"hipaa":true,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.6,"diarization_streaming":true,"turn_detection":null,"keyterms":true,"pii_redaction":true,"interim_results":true,"code_switching":null,"telephony_8k":false,"bills_silence":true,"_notes":"$0.01/min us-east-1 price list (older articles quote $0.024/min). 77 languages support streaming, some not in every region. Free tier 60 min/month for 12 months. Accepts 8 kHz PCM, FLAC or Opus but not mulaw. AWS asks you to send silence rather than gaps, so silence is billed. Also HTTP/2 transport. EU via AWS EU regions.","_added":["languages: https://docs.aws.amazon.com/transcribe/latest/dg/supported-languages.html"]}},{"id":"elevenlabs-scribe-realtime","name":"ElevenLabs Scribe v2 Realtime","vendor":"ElevenLabs","category":"stt","summary":"WebSocket realtime STT (scribe_v2_realtime) with VAD or manual commits, 90+ languages, keyterms and entity detection, billed from your ElevenLabs plan credits.","status":"GA","models":[{"name":"scribe_v2_realtime","status":"GA (only realtime model)","notes":"Batch counterpart is Scribe v2 (has diarization); realtime does not list diarization."}],"transports":["WebSocket"],"audio":{"input":"audio_format: pcm_8000, pcm_16000 (default), pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000; audio sent base64 in JSON input_audio_chunk messages.","output":"n/a"},"languages":"90+ languages (vendor); language_code plus secondary_languages hints.","latency":"Vendor claim: ~150 ms (footnoted, conditions not stated).","features":["partial_transcript and committed_transcript","commit_strategy: manual, vad, or turn_prediction","VAD tuning (threshold, silence secs, min speech/silence ms)","word timestamps (include_timestamps)","language detection","keyterms (up to 50, 20 chars each; paid)","entity detection","experimental transcript editing (paid)","background audio filter","enable_logging=false option","regional and data-residency hosts"],"pricing":{"model":"per-hour","items":[{"what":"Scribe v2 Realtime","price":"$0.39","unit":"per hour","notes":"Same rate on every plan; included hours scale with plan (Free 2h30, Starter 15h, Creator 56h, Pro 254h, Scale 767h, Business 2,538h)."},{"what":"Keyterm prompting (realtime)","price":"+$0.08","unit":"per hour","notes":""},{"what":"Transcript editing (realtime)","price":"+$0.12","unit":"per hour","notes":"Experimental."},{"what":"Scribe v2 batch (reference)","price":"$0.22","unit":"per hour","notes":""}],"est_per_minute_usd":{"low":0.0065,"high":0.0098,"basis":"$0.39/hr base to $0.59/hr with keyterms + editing."},"free_tier":"Free plan includes about 2.5 realtime hours.","source":"https://elevenlabs.io/pricing/api"},"limits":["Realtime concurrency limit by plan is in a separate chart on the models page (not captured).","keepalive_interval_ms configurable 500-10000; errors include session_time_limit_exceeded, commit_throttled, queue_overflow, insufficient_audio_activity."],"regions":"Default api.elevenlabs.io; US host and EU, India, Singapore data-residency hosts listed.","setup":{"steps":["Create an API key.","Server: connect to wss://api.elevenlabs.io/v1/speech-to-text/realtime?model_id=scribe_v2_realtime&audio_format=pcm_16000&commit_strategy=vad with header xi-api-key.","Browser: create a single-use token via the tokens endpoint on your server and pass ?token=.","Send input_audio_chunk JSON messages with base64 audio; read partial_transcript and committed_transcript."],"endpoint":"wss://api.elevenlabs.io/v1/speech-to-text/realtime","auth":"xi-api-key header, or single-use token in the token query parameter for client-side use.","snippet_lang":"python","snippet":"import asyncio, base64, json, os, websockets  # pip install websockets>=14\n\nURL = (\"wss://api.elevenlabs.io/v1/speech-to-text/realtime\"\n       \"?model_id=scribe_v2_realtime&audio_format=pcm_16000&commit_strategy=vad\")\n\nasync def main():\n    hdr = {\"xi-api-key\": os.environ[\"ELEVENLABS_API_KEY\"]}\n    async with websockets.connect(URL, additional_headers=hdr) as ws:\n        async def send():\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:\n                while chunk := f.read(3200):  # 100 ms\n                    await ws.send(json.dumps({\n                        \"message_type\": \"input_audio_chunk\",\n                        \"audio_base_64\": base64.b64encode(chunk).decode(),\n                        \"commit\": False, \"sample_rate\": 16000}))\n                    await asyncio.sleep(0.1)\n        asyncio.create_task(send())\n        async for msg in ws:\n            d = json.loads(msg)\n            t = d.get(\"message_type\")\n            if t == \"partial_transcript\":\n                print(\"partial\", d.get(\"text\"))\n            elif t == \"committed_transcript\":\n                print(\"FINAL\", d.get(\"text\"))\n            elif t and t.endswith(\"error\"):\n                print(d); break\n\nasyncio.run(main())"},"warnings":[{"severity":"medium","title":"No realtime diarization listed","detail":"Speaker diarization is documented for batch Scribe v2 only. For multi-party calls, send separate channels as separate sessions."},{"severity":"medium","title":"Logging on by default","detail":"enable_logging defaults to true; set it to false if you need no retention (may require an eligible plan)."},{"severity":"medium","title":"Base64 JSON overhead","detail":"Audio goes as base64 inside JSON (about 33% larger than binary frames), which matters on mobile uplinks."},{"severity":"low","title":"Plan credits, not pure PAYG","detail":"Realtime STT draws from subscription credits; included hours depend on the plan, and overage behaviour follows your plan settings."},{"severity":"low","title":"Keyterm limits are small","detail":"Realtime supports at most 50 keyterms of 20 characters each, and they cost +$0.08/hr."}],"best_for":"Teams already using ElevenLabs TTS/agents who want the same vendor for live STT, including 8 kHz mu-law telephony.","open_source":false,"self_hostable":false,"compliance":"Data-residency hosts for EU, India, Singapore; enable_logging=false for zero retention. Certifications not re-verified.","docs":[{"label":"Realtime API reference","url":"https://elevenlabs.io/docs/api-reference/speech-to-text/v-1-speech-to-text-realtime"},{"label":"STT capability page","url":"https://elevenlabs.io/docs/capabilities/speech-to-text"},{"label":"API pricing","url":"https://elevenlabs.io/pricing/api"}],"sources":["https://elevenlabs.io/pricing/api","https://elevenlabs.io/docs/capabilities/speech-to-text","https://elevenlabs.io/docs/api-reference/speech-to-text/v-1-speech-to-text-realtime"],"confidence":"medium","unverified":"Realtime concurrency per plan, whether enable_logging=false is limited to enterprise, exact error message_type names used in the snippet's error check, and the meaning of the ~150 ms footnote.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"ElevenLabs Scribe","facts":{"cat":"stt","latency_ms":150,"languages":90,"max_session_min":null,"concurrency":9,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.39,"diarization_streaming":null,"turn_detection":true,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":null,"telephony_8k":true,"bills_silence":null,"_notes":"~150 ms vendor claim, conditions not stated. 90+ languages. Concurrency 9 is the Starter plan; Pro 30, Scale 45, Free 6. Free plan includes about 2.5 realtime hours. Diarization not listed for realtime. Keyterms +$0.08/hr. EU, India and Singapore residency hosts.","_added":["concurrency: https://elevenlabs.io/docs/models"]}},{"id":"mistral-voxtral-realtime","name":"Mistral Voxtral Mini Transcribe Realtime","vendor":"Mistral AI","category":"stt","summary":"Hosted realtime transcription over WebSocket at $0.006/min, and the same 4B model released as Apache-2.0 open weights for self-hosting.","status":"GA","models":[{"name":"voxtral-mini-transcribe-realtime-2602","status":"GA (API)","notes":"Only realtime-capable model; not compatible with the diarize parameter."},{"name":"mistralai/Voxtral-Mini-4B-Realtime-2602","status":"Open weights","notes":"Apache 2.0 on Hugging Face; 4B parameters."}],"transports":["WebSocket"],"audio":{"input":"AudioFormat(encoding='pcm_s16le', sample_rate=16000) in docs examples; mic example sends 480 ms chunks.","output":"n/a"},"languages":"13 languages (vendor).","latency":"Vendor claim: configurable latency down to sub-200 ms; press reports ~1-2% error rate at 480 ms delay (secondary source).","features":["streaming text deltas","configurable latency","short-lived rt_ tokens for browsers","open weights for on-prem"],"pricing":{"model":"per-minute","items":[{"what":"Voxtral Realtime API","price":"$0.006","unit":"per minute","notes":"From Mistral's Voxtral Transcribe 2 announcement; batch Voxtral is $0.003/min."},{"what":"Self-hosted open weights","price":"free","unit":"","notes":"Apache 2.0; you pay for GPUs."}],"est_per_minute_usd":{"low":0.006,"high":0.006,"basis":"Flat announced rate; not confirmed on a live pricing table in this pass."},"free_tier":"Not stated.","source":"https://mistral.ai/news/voxtral-transcribe-2"},"limits":["rt_ client tokens expire after about 900 seconds; mint them just before connecting.","Python realtime needs SDK v2 (mistralai>=2, extra [realtime]); not available in v1 SDK."],"regions":"Mistral API (EU company); self-host anywhere.","setup":{"steps":["pip install \"mistralai[realtime]>=2\" (V2 SDK).","Stream PCM to client.audio.realtime.transcribe_stream with model voxtral-mini-transcribe-realtime-2602.","For browsers: server calls POST https://api.mistral.ai/v1/client/sessions {purpose:'realtime', model} to mint an rt_ token; browser opens the WebSocket with subprotocols ['realtime', token].","Self-host: download the HF weights and serve with an engine that supports the realtime model (check model card)."],"endpoint":"wss://api.mistral.ai/v1/audio/transcriptions/realtime?model=voxtral-mini-transcribe-realtime-2602","auth":"Authorization: Bearer <key> server side; rt_ short-lived token via Sec-WebSocket-Protocol in browsers.","snippet_lang":"python","snippet":"import asyncio, os  # pip install \"mistralai[realtime]>=2\"\nfrom mistralai.client import Mistral\nfrom mistralai.client.models import (AudioFormat, TranscriptionStreamDone,\n                                     TranscriptionStreamTextDelta)\n\nclient = Mistral(api_key=os.environ[\"MISTRAL_API_KEY\"])\nfmt = AudioFormat(encoding=\"pcm_s16le\", sample_rate=16000)\n\nasync def audio_chunks():\n    with open(\"audio_16k_mono.raw\", \"rb\") as f:\n        while chunk := f.read(15360):  # 480 ms\n            yield chunk\n            await asyncio.sleep(0.48)\n\nasync def main():\n    async for ev in client.audio.realtime.transcribe_stream(\n        audio_stream=audio_chunks(),\n        model=\"voxtral-mini-transcribe-realtime-2602\",\n        audio_format=fmt,\n    ):\n        if isinstance(ev, TranscriptionStreamTextDelta):\n            print(ev.text, end=\"\", flush=True)\n        elif isinstance(ev, TranscriptionStreamDone):\n            print(\"\\n[done]\")\n\nasyncio.run(main())"},"warnings":[{"severity":"medium","title":"No diarization in realtime","detail":"Mistral's docs state realtime is not compatible with the diarize parameter; use one or the other."},{"severity":"medium","title":"Only 13 languages","detail":"Strong performance is claimed for 13 languages, far fewer than Soniox, Gladia or Google. Check your language first."},{"severity":"medium","title":"SDK v1 cannot do realtime","detail":"Realtime transcription is only in the V2 Python SDK (mistralai>=2); v1 code examples will not work."},{"severity":"low","title":"Short token lifetime","detail":"Browser rt_ tokens last about 900 s; mint right before the user starts talking, not at page load."},{"severity":"low","title":"Price from announcement","detail":"The $0.006/min rate comes from the launch post; confirm on the console pricing table before committing."}],"best_for":"EU-vendor realtime STT with an escape hatch to run the identical model on your own GPUs.","open_source":true,"self_hostable":true,"hardware":"4B-parameter model; GPU recommended for realtime serving. Mistral says the footprint can run on edge devices (not benchmarked here).","license":"Apache 2.0 (open weights).","compliance":"Not re-verified. Self-hosting keeps audio fully in your infrastructure.","docs":[{"label":"Realtime transcription docs","url":"https://docs.mistral.ai/studio/audio/speech_to_text/realtime_transcription"},{"label":"Client authentication","url":"https://docs.mistral.ai/studio/audio/speech_to_text/realtime_transcription/client_auth"},{"label":"Announcement","url":"https://mistral.ai/news/voxtral-transcribe-2"},{"label":"Open weights","url":"https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602"}],"sources":["https://docs.mistral.ai/studio/audio/speech_to_text/realtime_transcription.md","https://docs.mistral.ai/studio/audio/speech_to_text/realtime_transcription/client_auth.md","https://docs.mistral.ai/capabilities/audio_transcription","https://mistral.ai/news/voxtral-transcribe-2","https://the-decoder.com/voxtral-transcribe-2-offers-speech-recognition-at-0-003-per-minute/","https://huggingface.co/api/models/mistralai/Voxtral-Mini-4B-Realtime-2602"],"confidence":"medium","unverified":"Price on a live pricing table, language list, concurrency limits, and the recommended self-host serving stack.","cat":"stt","kind":"stt","verified_at":"2026-10-10","short":"Voxtral Realtime","facts":{"cat":"stt","latency_ms":200,"languages":13,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":true,"price_per_hour":0.36,"diarization_streaming":false,"turn_detection":null,"keyterms":null,"pii_redaction":null,"interim_results":true,"code_switching":null,"telephony_8k":null,"bills_silence":null,"_notes":"Latency configurable down to sub-200 ms (vendor). $0.006/min from the launch announcement, not a live price table. Open weights Apache 2.0 (Voxtral-Mini-4B-Realtime). Realtime cannot be combined with diarization."}},{"id":"cartesia-ink","name":"Cartesia Ink (ink-2, ink-whisper)","vendor":"Cartesia","category":"stt","summary":"Streaming STT over WebSocket aimed at voice agents, billed in plan credits per second of audio (silence included). ink-2 is the newer model; ink-whisper is the cheaper Whisper-based one.","status":"GA","models":[{"name":"ink-2","status":"GA","notes":"3 credits per second; keyterms supported (up to 100 terms / 1200 chars)."},{"name":"ink-whisper","status":"GA","notes":"1 credit per second streaming; language, min_volume and max_silence_duration_secs params."},{"name":"ink-preview","status":"Preview","notes":"Accepted model value on the WebSocket."}],"transports":["WebSocket"],"audio":{"input":"encoding pcm_s16le, pcm_s32le, pcm_f16le, pcm_f32le, pcm_mulaw, pcm_alaw with explicit sample_rate; ~100 ms binary chunks.","output":"n/a"},"languages":"ink-whisper takes an ISO-639-1 language (default en); ink-2 language coverage not captured.","latency":"No numeric claim captured.","features":["interim and final transcripts (is_final)","finalize command to flush","word timestamps","keyterms (ink-2)","turn-based endpoint /stt/turns/websocket","access tokens for browsers"],"pricing":{"model":"per-hour","items":[{"what":"ink-2 streaming","price":"3 credits","unit":"per second of audio","notes":"Vendor: about $0.39/hr on the Scale plan."},{"what":"ink-whisper streaming","price":"1 credit","unit":"per second of audio","notes":"Vendor: about $0.13/hr on Scale. Batch /stt is 1 credit per 2 s."},{"what":"Plans","price":"Free $0 (20K credits), Pro $5 (100K), Startup $49 (1.25M), Scale $299 (8M)","unit":"per month","notes":"Overage $65 / $45 / $38 per 1M credits on Pro / Startup / Scale."}],"est_per_minute_usd":{"low":0.0022,"high":0.0117,"basis":"ink-whisper on Scale ($0.13/hr) up to ink-2 on Pro overage (10,800 credits/hr x $65/1M = $0.70/hr)."},"free_tier":"Free plan: 20K credits/month (about 1h51m of ink-2 per plan table).","source":"https://docs.cartesia.ai/pricing"},"limits":["STT concurrent requests: Free 8, Pro 12, Startup 20, Scale 60, Enterprise custom (separate pool from TTS).","With overages disabled, requests fail when credits run out."],"regions":"Not captured.","setup":{"steps":["Create an API key at play.cartesia.ai.","Connect to wss://api.cartesia.ai/stt/websocket?model=ink-2&encoding=pcm_s16le&sample_rate=16000&cartesia_version=2026-08-14 with header X-API-Key.","Send binary audio; send text 'finalize' to flush, 'close' to end.","Browsers: mint a short-lived access token server-side and pass access_token= in the URL."],"endpoint":"wss://api.cartesia.ai/stt/websocket (turn-based: /stt/turns/websocket)","auth":"X-API-Key header server-side; access_token query param for clients.","snippet_lang":"python","snippet":"import asyncio, json, os, websockets  # pip install websockets>=14\n\nURL = (\"wss://api.cartesia.ai/stt/websocket?model=ink-2&encoding=pcm_s16le\"\n       \"&sample_rate=16000&cartesia_version=2026-08-14\")\n\nasync def main():\n    hdr = {\"X-API-Key\": os.environ[\"CARTESIA_API_KEY\"]}\n    async with websockets.connect(URL, additional_headers=hdr) as ws:\n        async def send():\n            with open(\"audio_16k_mono.raw\", \"rb\") as f:\n                while chunk := f.read(3200):  # 100 ms\n                    await ws.send(chunk)\n                    await asyncio.sleep(0.1)\n            await ws.send(\"finalize\")\n            await ws.send(\"close\")\n        asyncio.create_task(send())\n        final = \"\"\n        async for msg in ws:\n            d = json.loads(msg)\n            if d[\"type\"] == \"transcript\":\n                if d[\"is_final\"]:\n                    final += d[\"text\"]  # deltas: do not strip or add spaces\n                print(\"FINAL\" if d[\"is_final\"] else \"partial\", d[\"text\"])\n            elif d[\"type\"] in (\"done\", \"error\"):\n                print(final or d); break\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Silence is billed","detail":"Cartesia's pricing docs state STT charges include silence even when no transcript is produced. Do not leave sockets streaming idle audio."},{"severity":"medium","title":"Credits shared with TTS","detail":"STT and TTS draw from the same monthly credit pool; heavy TTS usage can starve STT and vice versa. Turn on overages or monitor balance."},{"severity":"medium","title":"ink-2 costs 3x ink-whisper","detail":"ink-2 is 3 credits/s vs 1 credit/s for ink-whisper; pick per use case."},{"severity":"medium","title":"Text is a delta","detail":"Docs warn not to strip whitespace or insert spaces between chunks; concatenate is_final chunks exactly as received."},{"severity":"low","title":"Version pinning required","detail":"cartesia_version is a required query parameter (2026-08-14 in docs); old versions may change behaviour."}],"best_for":"Voice agents already using Cartesia Sonic TTS who want one credit pool and turn-aware STT.","open_source":false,"self_hostable":false,"compliance":"Enterprise plans list DPAs and BAAs; not re-verified.","docs":[{"label":"STT WebSocket reference","url":"https://docs.cartesia.ai/api-reference/stt/stt"},{"label":"Pricing docs","url":"https://docs.cartesia.ai/pricing"},{"label":"Plans","url":"https://cartesia.ai/pricing"}],"sources":["https://docs.cartesia.ai/api-reference/stt/stt","https://docs.cartesia.ai/pricing.md","https://cartesia.ai/pricing","https://cartesia.ai/ink"],"confidence":"medium","unverified":"ink-2 language list and latency; Free-plan STT concurrency of 8 read from a plan-comparison table that may be ordered differently.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":12,"free_tier":true,"free_credit_usd":null,"hipaa":true,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.39,"diarization_streaming":null,"turn_detection":true,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":null,"telephony_8k":true,"bills_silence":true,"_notes":"Credit-based. $0.39/hr is ink-2 on the Scale plan (vendor); ink-2 on Pro overage is about $0.70/hr; ink-whisper about $0.13/hr on Scale. Concurrency 12 is the Pro plan (Free 8, Scale 60). Free plan 20K credits/month shared with TTS. BAA on Enterprise plans only. Turn detection via the /stt/turns endpoint."}},{"id":"smallest-pulse","name":"Smallest.ai Pulse STT","vendor":"Smallest AI","category":"stt","summary":"Unified STT endpoint with Pulse (multilingual, streaming) and Pulse 2.0 (English streaming with built-in end-of-turn, emotion, gender). Strong Indic and East Asian coverage, region-dependent.","status":"GA","models":[{"name":"pulse","status":"GA","notes":"Multilingual, streaming and pre-recorded."},{"name":"pulse-2","status":"GA","notes":"English only, streaming only, end-of-turn detection plus per-utterance emotion and gender."},{"name":"pulse-pro","status":"GA (HTTP only)","notes":"Rejected on the streaming endpoint with 400 before upgrade."}],"transports":["WebSocket"],"audio":{"input":"See realtime quickstart (formats not captured in this pass).","output":"n/a"},"languages":"Streaming: en, hi, de, es, ru, it, fr, nl, pt, zh, yue, ja, ko, gu, mr, or, bn, ta, te, kn, ml plus regional auto-detect groups (north_indic, multi-asian, multi-south-indic). East Asian languages US region only; South Indian languages India region only (beta).","latency":"See Pulse model card (not captured).","features":["streaming partials","end-of-turn detection (pulse-2)","diarization","word and sentence timestamps","emotion and gender detection","PII and PCI redaction","regional auto language detection"],"pricing":{"model":"per-minute","items":[{"what":"Pulse streaming (WebSocket)","price":"$0.006","unit":"per minute","notes":"Standard plan, per 2026-05-28 changelog. An older pricing page showed ~$0.004/min."},{"what":"Pulse non-streaming (reference)","price":"$0.0035","unit":"per minute","notes":""}],"est_per_minute_usd":{"low":0.004,"high":0.006,"basis":"Changelog $0.006/min is the newest first-party figure; older page $0.004/min."},"free_tier":"Not captured.","source":"https://docs.smallest.ai/waves/changelog/pulse-stt/2026/5/28"},"limits":["Language availability depends on region host (api.us.smallest.ai vs api.smallest.ai); wrong region returns LANGUAGE_NOT_ENABLED_IN_REGION."],"regions":"US (api.us.smallest.ai) and India (api.smallest.ai) hosts.","setup":{"steps":["Get an API key from the Smallest console.","Open WS /waves/v1/stt/live?model=pulse (or pulse-2) on the correct regional host.","Follow the realtime quickstart for the audio format and message schema."],"endpoint":"wss://api.smallest.ai/waves/v1/stt/live?model=pulse (US: wss://api.us.smallest.ai/waves/v1/stt/live)","auth":"API key (header format not captured; see quickstart).","snippet_lang":"python","snippet":""},"warnings":[{"severity":"medium","title":"Conflicting published prices","detail":"Smallest's changelog lists $0.006/min for streaming Pulse while an older pricing page shows ~$0.004/min. Confirm in your dashboard."},{"severity":"medium","title":"Region decides languages","detail":"East Asian streaming languages are US-only and South Indian ones India-only; connecting to the wrong host is rejected."},{"severity":"medium","title":"Pulse Pro is not streamable","detail":"The most accurate English model (pulse-pro) is HTTP-only; streaming requests are rejected with 400."},{"severity":"low","title":"pulse-2 is English only","detail":"Built-in end-of-turn detection exists only on the English pulse-2 model."}],"best_for":"Indian-language and Asian-language voice agents that want turn detection and emotion signals.","open_source":false,"self_hostable":false,"compliance":"Zero data retention documented for Enterprise plans.","docs":[{"label":"STT overview","url":"https://docs.smallest.ai/models/speech-to-text/overview"},{"label":"Changelog (pricing)","url":"https://docs.smallest.ai/waves/changelog/pulse-stt/2026/5/28"}],"sources":["https://docs.smallest.ai/models/speech-to-text/overview.md","https://docs.smallest.ai/waves/changelog/pulse-stt/2026/5/28","https://smallest.ai/pricing-page-2"],"confidence":"low","unverified":"Audio formats, auth header, message schema, latency, concurrency and current price. No snippet given because the protocol was not verified.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":21,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":false,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.36,"diarization_streaming":true,"turn_detection":true,"keyterms":null,"pii_redaction":true,"interim_results":true,"code_switching":null,"telephony_8k":null,"bills_silence":null,"_notes":"$0.006/min from the May 2026 changelog; an older page showed ~$0.004/min. 21 streaming languages, availability depends on US or India host. Hosts are US and India only. End-of-turn only on English pulse-2. Diarization and PII redaction are listed features; live support not confirmed."}},{"id":"sarvam-saaras","name":"Sarvam AI Saaras Realtime STT","vendor":"Sarvam AI","category":"stt","summary":"Indian-language focused STT with a 2026 Realtime WebSocket API (saaras:v3-realtime default, saaras:v4) priced in rupees per hour.","status":"GA","models":[{"name":"saaras:v3-realtime","status":"GA (default on Realtime API)","notes":"Introduced Aug 2026 for voice agents and live transcription."},{"name":"saaras:v4","status":"GA","notes":"Available on REST, WebSocket, batch and the Realtime API (as saaras:v4 / v4-realtime per changelog)."},{"name":"saaras:v3 legacy streaming WebSocket","status":"Legacy","notes":"Older streaming socket, still available."}],"transports":["WebSocket"],"audio":{"input":"encoding and sample_rate parameters on the realtime channel (values not captured).","output":"n/a"},"languages":"Indian languages plus English (exact list not captured).","latency":"Not captured.","features":["realtime transcription","speech-to-text with translation","diarization (priced variant)"],"pricing":{"model":"per-hour","items":[{"what":"Speech to Text","price":"INR 30","unit":"per hour, billed per second","notes":"Pricing table does not separate streaming from batch."},{"what":"STT with diarization","price":"INR 45","unit":"per hour","notes":""},{"what":"STT and translate","price":"INR 30","unit":"per hour","notes":""}],"est_per_minute_usd":{"low":0.0059,"high":0.0088,"basis":"INR 30-45/hr converted at about INR 85 per USD (third-party FX note, July 2026); currency risk applies."},"free_tier":"INR 100 credits for new users.","source":"https://docs.sarvam.ai/api/pricing"},"limits":["See Credits and Rate Limits page (not captured)."],"regions":"India (not re-verified).","setup":{"steps":["Get an API subscription key from the Sarvam dashboard.","Connect to wss://api.sarvam.ai/speech-to-text-realtime/ws with model, language_code, encoding and sample_rate parameters.","Server: send header Api-Subscription-Key; browser: use WebSocket subprotocol 'api-subscription-key.<key>' (exposes the key, so prefer a server relay)."],"endpoint":"wss://api.sarvam.ai/speech-to-text-realtime/ws","auth":"Api-Subscription-Key header, or subprotocol api-subscription-key.<key> in browsers.","snippet_lang":"python","snippet":""},"warnings":[{"severity":"medium","title":"Browser auth exposes the key","detail":"The documented browser method puts the subscription key in the WebSocket subprotocol; no short-lived token mechanism was found. Relay through your server."},{"severity":"medium","title":"Two streaming APIs","detail":"The newer Realtime API (saaras:v3-realtime) supersedes the legacy streaming WebSocket for new work; message schemas differ, so do not mix examples."},{"severity":"low","title":"Rupee pricing","detail":"Prices are published only in INR; USD cost moves with the exchange rate."},{"severity":"low","title":"Diarization raises price 50%","detail":"STT with diarization is INR 45/hr vs INR 30/hr."}],"best_for":"Voice agents and transcription for Indian languages and Hinglish.","open_source":false,"self_hostable":false,"compliance":"Not re-verified.","docs":[{"label":"Pricing","url":"https://docs.sarvam.ai/api/pricing"},{"label":"Realtime STT WebSocket","url":"https://docs.sarvam.ai/api-reference/speech-to-text/transcribe/realtime/ws"},{"label":"Changelog Aug 2026","url":"https://docs.sarvam.ai/changelog/2026/8/1"}],"sources":["https://docs.sarvam.ai/api/pricing.md","https://docs.sarvam.ai/api-reference/speech-to-text/transcribe/realtime/ws.md","https://docs.sarvam.ai/changelog/2026/8/1","https://usagepricing.com/blueprint/sarvam-ai"],"confidence":"low","unverified":"Message schema (no snippet given), language list, rate limits, latency, and whether the realtime API has a distinct price.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":0.35,"diarization_streaming":null,"turn_detection":null,"keyterms":null,"pii_redaction":null,"interim_results":null,"code_switching":null,"telephony_8k":null,"bills_silence":null,"_notes":"Priced in INR: INR 30/hr (INR 45/hr with diarization), converted at about INR 85 per USD. Free credit INR 100. Indian languages plus English, exact list not captured. Pricing does not separate streaming from batch."}},{"id":"groq-whisper","name":"Groq Whisper (no streaming)","vendor":"Groq","category":"stt","summary":"Extremely fast but file-based Whisper transcription. There is no streaming/WebSocket STT endpoint; 'realtime' use means chunking audio yourself and paying a 10-second minimum per request.","status":"GA","models":[{"name":"whisper-large-v3-turbo","status":"GA","notes":"$0.04/hr, transcription only (no translation), 216x real-time speed factor (vendor)."},{"name":"whisper-large-v3","status":"GA","notes":"$0.111/hr, transcription and translation."}],"transports":["HTTPS (OpenAI-compatible, file upload)"],"audio":{"input":"flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm files or URL; 25 MB free tier / 100 MB dev tier.","output":"n/a"},"languages":"Multilingual (Whisper).","latency":"Vendor: real-time speed factor 189x (v3) to 216x (turbo) for files; not a streaming latency.","features":["segment and word timestamps","translation to English (large-v3 only)","OpenAI-compatible endpoint"],"pricing":{"model":"per-hour","items":[{"what":"whisper-large-v3-turbo","price":"$0.04","unit":"per hour","notes":"Minimum billed length 10 s per request."},{"what":"whisper-large-v3","price":"$0.111","unit":"per hour","notes":"Minimum billed length 10 s per request."}],"est_per_minute_usd":{"low":0.00067,"high":0.0019,"basis":"File pricing; chunking live audio into short requests raises effective cost because each request bills at least 10 s."},"free_tier":"Groq free tier with rate limits (not detailed here).","source":"https://console.groq.com/docs/speech-to-text"},"limits":["No streaming endpoint.","10-second minimum billed length per request.","Free tier file size 25 MB."],"regions":"Groq cloud (not detailed).","setup":{"steps":["Get a Groq API key.","POST audio to https://api.groq.com/openai/v1/audio/transcriptions (OpenAI SDK compatible).","For pseudo-live use, cut audio on VAD pauses and send each utterance as a file."],"endpoint":"https://api.groq.com/openai/v1/audio/transcriptions","auth":"Authorization: Bearer <GROQ_API_KEY>","snippet_lang":"python","snippet":"import os\nfrom openai import OpenAI  # pip install openai\n\nclient = OpenAI(api_key=os.environ[\"GROQ_API_KEY\"], base_url=\"https://api.groq.com/openai/v1\")\nwith open(\"utterance.wav\", \"rb\") as f:  # one VAD-segmented utterance, not a live stream\n    r = client.audio.transcriptions.create(model=\"whisper-large-v3-turbo\", file=f)\nprint(r.text)"},"warnings":[{"severity":"high","title":"Not a streaming API","detail":"Groq has only file transcription and translation endpoints. There are no partial results; you must segment audio yourself and accept utterance-level latency."},{"severity":"medium","title":"10-second billing floor","detail":"Requests shorter than 10 s are billed as 10 s, so chopping speech into 1-2 s chunks can multiply cost 5-10x."},{"severity":"medium","title":"Whisper hallucinations on silence","detail":"Sending silent or noise-only chunks to Whisper models can produce invented text; gate requests with a VAD."},{"severity":"low","title":"Turbo cannot translate","detail":"whisper-large-v3-turbo supports transcription only; translation needs whisper-large-v3."}],"best_for":"Cheap near-real-time transcription of short VAD-segmented utterances where partials are not needed.","open_source":false,"self_hostable":false,"compliance":"Not re-verified.","docs":[{"label":"Speech to text","url":"https://console.groq.com/docs/speech-to-text"}],"sources":["https://console.groq.com/docs/speech-to-text.md"],"confidence":"high","unverified":"Rate limits and compliance.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":false,"open_weights":true,"price_per_hour":0.04,"diarization_streaming":false,"turn_detection":false,"keyterms":null,"pii_redaction":null,"interim_results":false,"code_switching":null,"telephony_8k":null,"bills_silence":null,"_notes":"Not a streaming API: file upload over HTTPS only. $0.04/hr is whisper-large-v3-turbo; large-v3 is $0.111/hr. 10 s minimum billed per request, so short chunks cost more. Open weights refers to the Whisper models; the Groq service is hosted only."}},{"id":"fireworks-streaming-asr","name":"Fireworks AI streaming ASR (deprecated)","vendor":"Fireworks AI","category":"stt","summary":"Fireworks offered a WebSocket streaming transcription API (v1 and a v2 preview), but its changelog marks audio inference as deprecated on 2026-06-10 and the ASR docs page now returns not found.","status":"Deprecated","models":[{"name":"streaming-speech (v1)","status":"Deprecated","notes":"Launched at $0.0032/audio min (launch blog)."},{"name":"fireworks-asr-v2 / streaming v2","status":"Deprecated (was preview)","notes":"Blog listed $0.0035/audio min."}],"transports":["WebSocket"],"audio":{"input":"n/a (deprecated)","output":"n/a"},"languages":"n/a","latency":"n/a","features":[],"pricing":{"model":"per-minute","items":[{"what":"Historical streaming v2","price":"$0.0035","unit":"per audio minute","notes":"From a 2025 blog; not current."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Do not plan on this service."},"free_tier":"n/a","source":"https://docs.fireworks.ai/updates/changelog"},"limits":["Service deprecated."],"regions":"n/a","setup":{"steps":["Do not start new integrations; migrate existing ones to another provider."],"endpoint":"Formerly wss://audio-streaming.us-virginia-1.direct.fireworks.ai/v1/audio/transcriptions/streaming","auth":"n/a","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"Audio inference deprecated","detail":"Fireworks' changelog entry dated 2026-06-10 says audio inference and image generation are deprecated, and the querying-asr-models docs page is gone. Existing integrations should migrate."},{"severity":"medium","title":"Stale third-party listings","detail":"Model catalogs and comparison sites still list Fireworks streaming ASR prices; treat them as outdated."},{"severity":"low","title":"Audio input to LLMs is different","detail":"Fireworks still supports audio as input to multimodal chat models (e.g. Qwen3 Omni), which is not a streaming ASR replacement."}],"best_for":"Nothing new; listed so readers know it is gone.","open_source":false,"self_hostable":false,"compliance":"n/a","docs":[{"label":"Changelog","url":"https://docs.fireworks.ai/updates/changelog"}],"sources":["https://docs.fireworks.ai/updates/changelog.md","https://fireworks.ai/blog/audio-september-release","https://fireworks.ai/blog/streaming-audio-launch"],"confidence":"medium","unverified":"Exact shutdown date for existing customers.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_hour":null,"diarization_streaming":null,"turn_detection":null,"keyterms":null,"pii_redaction":null,"interim_results":null,"code_switching":null,"telephony_8k":null,"bills_silence":null,"_notes":"Deprecated: audio inference deprecated in the 2026-06-10 changelog. Historical price $0.0035/min ($0.21/hr)."}},{"id":"picovoice-cheetah","name":"Picovoice Cheetah (on-device streaming STT)","vendor":"Picovoice","category":"stt","summary":"Commercial on-device streaming STT SDK running locally on desktop, mobile, web and Raspberry Pi. Leopard is the on-device batch counterpart. Requires a Picovoice AccessKey.","status":"GA","models":[{"name":"Cheetah","status":"GA","notes":"Streaming, on-device, custom vocabulary and keyword boosting via Picovoice Console."},{"name":"Leopard","status":"GA","notes":"On-device batch (file) transcription; not streaming."}],"transports":["On-device SDK (no network audio)"],"audio":{"input":"16-bit PCM frames of Cheetah's frame_length (16 kHz mono in Picovoice SDKs; not restated on the doc page read).","output":"n/a"},"languages":"English, French, German, Italian, Japanese, Korean, Spanish, Portuguese; more for Enterprise.","latency":"Runs locally; no network round trip. No numeric claim captured.","features":["partial transcripts","endpoint detection (is_endpoint flag in SDK)","automatic punctuation (option)","custom vocabulary with optional IPA","keyword boosting"],"pricing":{"model":"free","items":[{"what":"Free tier","price":"$0","unit":"","notes":"Personal / non-commercial use; FAQ mentions up to 5 hours/month across listed engines. Third-party claim that the free tier ends 2026-06-30 is unconfirmed."},{"what":"Foundation Plan (startups)","price":"$6,000","unit":"per year","notes":"Startups under 5 years old and 20 employees or fewer; includes 25K Cheetah minutes/month plus other engines."},{"what":"Enterprise","price":"custom","unit":"","notes":""}],"est_per_minute_usd":{"low":0.02,"high":0.02,"basis":"Foundation Plan if fully used: $6,000 / (25,000 min x 12) = $0.02/min. Unused minutes raise the effective rate."},"free_tier":"Non-commercial free tier; commercial free trial grants Foundation Plan rights.","source":"https://picovoice.ai/checkout"},"limits":["Usage metered against account limits via AccessKey.","Commercial use requires a paid plan."],"regions":"On-device; audio never leaves the device.","setup":{"steps":["Sign up at Picovoice Console and copy your AccessKey.","pip install pvcheetah pvrecorder (SDKs also for Android, iOS, Web, C, .NET, etc.).","Feed frames of cheetah.frame_length samples; print partials; call flush() at endpoints."],"endpoint":"n/a (local library)","auth":"AccessKey passed to create(); validated against your account.","snippet_lang":"python","snippet":"import os\nimport pvcheetah, pvrecorder  # pip install pvcheetah pvrecorder\n\ncheetah = pvcheetah.create(access_key=os.environ[\"PV_ACCESS_KEY\"],\n                           enable_automatic_punctuation=True)\nrec = pvrecorder.PvRecorder(frame_length=cheetah.frame_length)\nrec.start()\ntry:\n    while True:\n        partial, is_endpoint = cheetah.process(rec.read())\n        print(partial, end=\"\", flush=True)\n        if is_endpoint:\n            print(cheetah.flush())\nexcept KeyboardInterrupt:\n    pass\nfinally:\n    rec.stop(); rec.delete(); cheetah.delete()"},"warnings":[{"severity":"high","title":"Free tier is non-commercial","detail":"Picovoice describes the free tier for personal non-commercial projects; shipping a product needs a paid plan (Foundation $6,000/yr for eligible startups, otherwise Enterprise)."},{"severity":"medium","title":"AccessKey still meters usage","detail":"Even though inference is local, the SDK requires an AccessKey that is checked against account limits; plan for that dependency and key protection in apps."},{"severity":"medium","title":"Limited languages","detail":"Eight languages out of the box; others only through Enterprise."},{"severity":"low","title":"Accuracy below cloud leaders","detail":"On-device models trade accuracy for privacy and cost; benchmark against your audio before choosing."}],"best_for":"Offline or privacy-critical apps (kiosks, embedded, mobile) that need streaming STT without sending audio to a cloud.","open_source":false,"self_hostable":true,"hardware":"Runs on CPU: Linux x86_64, macOS (x86_64, arm64), Windows (x86_64, arm64), Android, iOS, Web (WASM), Raspberry Pi 3/4/5.","license":"Proprietary SDK; AccessKey required.","compliance":"Audio processed on device (privacy by design).","docs":[{"label":"Cheetah docs","url":"https://picovoice.ai/docs/cheetah/"},{"label":"Foundation Plan checkout","url":"https://picovoice.ai/checkout"},{"label":"Free tier blog","url":"https://picovoice.ai/blog/introducing-picovoices-free-tier"}],"sources":["https://picovoice.ai/docs/cheetah/","https://picovoice.ai/checkout","https://picovoice.ai/blog/introducing-picovoices-free-tier","https://picovoice.ai/docs/faq/general"],"confidence":"medium","unverified":"Current free-tier status in 2026 (pricing page did not render), exact SDK signature details (create(enable_automatic_punctuation), process() returning (partial, is_endpoint)) are from prior SDK versions.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":8,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_hour":1.2,"diarization_streaming":null,"turn_detection":true,"keyterms":true,"pii_redaction":null,"interim_results":true,"code_switching":null,"telephony_8k":null,"bills_silence":null,"_notes":"On-device SDK, audio never leaves the device. Free tier is non-commercial only. Price is the Foundation Plan ($6,000/yr for 25K min/month) fully used; unused minutes raise the rate. 8 languages, more on Enterprise."}},{"id":"voxist","name":"Voxist ASR (via OVHcloud / Scaleway marketplaces)","vendor":"Voxist","category":"stt","summary":"French ASR vendor sold through OVHcloud AI Deploy and Scaleway marketplaces: a REST API with diarization and a WebSocket streaming API without diarization, licence billed per audio hour plus your GPU compute.","status":"GA","models":[{"name":"Voxist ASR (asr-rest:v1.0.9-gpu image for REST)","status":"GA","notes":"Streaming image and protocol in a separate guide."}],"transports":["WebSocket","REST"],"audio":{"input":"WAV, mono only; examples use 16 kHz.","output":"n/a"},"languages":"Not listed on the billing page (examples show en and French output).","latency":"Not captured.","features":["streaming recognition","punctuation","diarization (REST only)"],"pricing":{"model":"per-hour","items":[{"what":"Licence 0-100,000 min","price":"EUR 0.88","unit":"per hour of audio","notes":"Page examples are inconsistent about per-minute vs per-hour; confirm."},{"what":"Licence 100,001-500,000 min","price":"EUR 0.67","unit":"per hour of audio","notes":""},{"what":"Licence 3-6M min","price":"EUR 0.35","unit":"per hour of audio","notes":""},{"what":"GPU compute (OVHcloud AI Deploy)","price":"e.g. EUR 1.93","unit":"per GPU hour (V100S example)","notes":"Billed per second, separate from licence."}],"est_per_minute_usd":{"low":0.0063,"high":0.016,"basis":"Licence EUR 0.35-0.88/hr (~$0.006-0.016/min at ~1.1 USD/EUR) before GPU compute. Low confidence."},"free_tier":"None stated.","source":"https://docs.ovhcloud.com/en/guides/public-cloud/ai-machine-learning/ai-ecosystem-voxist-billing-features"},"limits":["WAV mono only.","Streaming has no diarization."],"regions":"Deployed in your OVHcloud/Scaleway project region (EU clouds); data location not stated on the page.","setup":{"steps":["Subscribe via OVHcloud AI Ecosystem or Scaleway marketplace.","Deploy the Voxist app on AI Deploy and wait for RUNNING.","Use <app_url>/docs for the API; streaming uses the WebSocket guide."],"endpoint":"<your AI Deploy app URL> (WebSocket path in Voxist streaming guide)","auth":"Authorization: Bearer <OVHcloud app token> for restricted apps","snippet_lang":"python","snippet":""},"warnings":[{"severity":"medium","title":"Two bills","detail":"You pay the Voxist licence per audio hour plus GPU compute per second for the deployed app, even while idle if the app keeps running."},{"severity":"medium","title":"Unit ambiguity","detail":"OVHcloud's page labels the table per hour of audio but its worked examples apply the rates per minute. Get a written quote."},{"severity":"low","title":"Strict input format","detail":"Only WAV mono is accepted; convert telephony mu-law or stereo first."}],"best_for":"EU sovereign-cloud deployments that want a French vendor running inside their own OVHcloud or Scaleway project.","open_source":false,"self_hostable":true,"compliance":"Runs in your own cloud project; certifications not checked.","docs":[{"label":"OVHcloud Voxist billing and features","url":"https://docs.ovhcloud.com/en/guides/public-cloud/ai-machine-learning/ai-ecosystem-voxist-billing-features"},{"label":"Scaleway marketplace","url":"https://www.scaleway.com/en/marketplace/catalog/voxist-speech-to-text/"}],"sources":["https://docs.ovhcloud.com/en/guides/public-cloud/ai-machine-learning/ai-ecosystem-voxist-billing-features","https://www.scaleway.com/en/marketplace/catalog/voxist-speech-to-text/"],"confidence":"low","unverified":"Streaming protocol, language list, latency and the exact price unit.","cat":"stt","kind":"stt","verified_at":"2026-10-10","facts":{"cat":"stt","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_hour":null,"diarization_streaming":false,"turn_detection":null,"keyterms":null,"pii_redaction":null,"interim_results":null,"code_switching":null,"telephony_8k":false,"bills_silence":null,"_notes":"Licence EUR 0.88/hr of audio (first 100K min) plus GPU compute billed separately; the price unit on the OVHcloud page is ambiguous. Runs in your own OVHcloud or Scaleway project (EU clouds). WAV mono only, so mulaw must be converted."}},{"id":"nvidia-nemotron-asr-streaming","name":"NVIDIA Nemotron ASR Streaming / Parakeet (Riva, NIM, NeMo)","vendor":"NVIDIA","category":"open-model","summary":"Open-weight cache-aware FastConformer-RNNT streaming ASR models (600M) that you serve with Riva/NIM containers or NeMo. Nemotron 3.5 ASR adds 40 language-locales in one checkpoint.","status":"GA","models":[{"name":"nvidia/nemotron-speech-streaming-en-0.6b (Nemotron 3 ASR)","status":"GA (open weights)","notes":"English, punctuation and capitalization; latency chosen at inference via att_context_size in 80 ms frames."},{"name":"nvidia/nemotron-3.5-asr-streaming-0.6b","status":"GA (open weights, HF 2026-06-04)","notes":"40 language-locales with language-ID prompt conditioning; GPU required."},{"name":"Parakeet CTC / RNNT / TDT family (e.g. parakeet-tdt-0.6b-v3)","status":"GA","notes":"Older Parakeet models; TDT v3 is CC-BY-4.0."},{"name":"nvidia/parakeet_realtime_eou_120m-v1","status":"Released 2025","notes":"Small realtime model with end-of-utterance detection."}],"transports":["gRPC (Riva/NIM)","WebSocket (Riva realtime client)","Python (NeMo)"],"audio":{"input":"16 kHz mono PCM typical for Riva streaming (not re-verified per model).","output":"n/a"},"languages":"English (Nemotron 3) or 40 language-locales (Nemotron 3.5).","latency":"Chunk sizes down to 80 ms per model card; NVIDIA FAQ reports 0.067 s ASR latency at 64 parallel streams for Parakeet CTC 1.1B on 3xH100 (vendor benchmark).","features":["cache-aware streaming (no re-processing overlap)","runtime latency/accuracy trade-off","punctuation and capitalization","batching many streams per GPU","fine-tuning with NeMo"],"pricing":{"model":"free","items":[{"what":"Open weights","price":"$0","unit":"","notes":"You pay for GPUs. NVIDIA AI Enterprise licence may apply to production Riva/NIM support."},{"what":"Hosted trial (build.nvidia.com)","price":"free trial credits","unit":"","notes":"Hosted nemotron-asr-streaming endpoint for evaluation."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Self-host compute only; cost depends on GPU price and streams per GPU."},"free_tier":"Open weights; hosted API trial on build.nvidia.com.","source":"https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b"},"limits":["GPU required for the multilingual model.","Third-party review notes occasional misplaced punctuation in streaming and inconsistent auto language detection; specify the language when known."],"regions":"Self-hosted anywhere.","setup":{"steps":["Pick a model on Hugging Face and read its licence (Nemotron EN: NVIDIA Open Model License; Nemotron 3.5: OpenMDW-1.1).","Production: run the Riva or NIM ASR container (NGC nemotron-asr-streaming) on an NVIDIA GPU.","Use nvidia-riva python-clients scripts (transcribe_mic.py, realtime_asr_client.py) to stream audio over gRPC.","Research/fine-tuning: use NeMo cache-aware streaming inference scripts."],"endpoint":"localhost:50051 gRPC (Riva/NIM default) or your service URL","auth":"None by default on self-hosted Riva; NGC API key to pull containers; NVIDIA API key for hosted trial.","snippet_lang":"bash","snippet":"# After starting a Riva/NIM ASR server on localhost:50051\npip install nvidia-riva-client\ngit clone https://github.com/nvidia-riva/python-clients && cd python-clients\npython scripts/asr/transcribe_mic.py --server localhost:50051 --language-code en-US\n# or stream a file at real-time pace:\npython scripts/asr/transcribe_file.py --server localhost:50051 --input-file audio_16k_mono.wav"},"warnings":[{"severity":"medium","title":"Licences differ per model","detail":"nemotron-speech-streaming-en uses the NVIDIA Open Model License, Nemotron 3.5 uses OpenMDW-1.1, Parakeet TDT v3 is CC-BY-4.0, and Riva/NIM containers fall under NVIDIA licence terms. Read each before shipping."},{"severity":"medium","title":"Multilingual model needs a GPU","detail":"Claims that the 40-language model runs on CPU refer to separate work on the English model; plan for NVIDIA GPUs."},{"severity":"medium","title":"You own scaling and uptime","detail":"Self-hosting means managing GPU capacity, batching, health checks and upgrades; throughput claims are vendor benchmarks on H100s."},{"severity":"low","title":"Conflicting release dates","detail":"Hugging Face shows Nemotron 3.5 released 2026-06-04 while an OpenRouter listing shows 2026-08-13 for a dated snapshot; pin the exact revision."}],"best_for":"High-volume, low-latency English or multilingual streaming on your own NVIDIA GPUs, especially for voice agents.","open_source":true,"self_hostable":true,"hardware":"NVIDIA GPU (H100/L40S/A10-class typical); many concurrent streams per GPU via batching.","license":"Nemotron 3 EN: NVIDIA Open Model License; Nemotron 3.5: OpenMDW-1.1; Parakeet TDT 0.6B v3: CC-BY-4.0; parakeet_realtime_eou: NVIDIA licence (other).","compliance":"Self-hosted: data stays in your environment.","docs":[{"label":"Nemotron 3.5 ASR streaming","url":"https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b"},{"label":"Nemotron ASR streaming EN","url":"https://huggingface.co/nvidia/nemotron-speech-streaming-en-0.6b"},{"label":"NGC container","url":"https://catalog.ngc.nvidia.com/orgs/nim/nvidia/containers/nemotron-asr-streaming/"},{"label":"Riva python clients","url":"https://github.com/nvidia-riva/python-clients"}],"sources":["https://huggingface.co/api/models/nvidia/nemotron-3.5-asr-streaming-0.6b","https://huggingface.co/nvidia/nemotron-speech-streaming-en-0.6b","https://perspectives.nvidia.com/nemotron-speech/task/faq/open-speech-models-proven-in-production/","https://pasqualepillitteri.it/en/news/5210/nemotron-3-5-asr-nvidia-40-languages-real-time","https://api.github.com/repos/nvidia-riva/python-clients/contents/scripts/asr"],"confidence":"medium","unverified":"Exact CLI flags of the Riva client scripts, Riva/NIM licence terms for production, and per-GPU stream counts.","cat":"open","kind":"stt","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":40,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":0.6,"vram_gb":null,"cpu_ok":false,"commercial_ok":true,"license_short":"OpenMDW-1.1","full_duplex":null,"streaming":true,"_notes":"Values for Nemotron 3.5 ASR (40 language-locales); the English Nemotron 3 model uses the NVIDIA Open Model License. Chunks down to 80 ms; no end-to-end latency figure for this model. Hosted trial on build.nvidia.com."}},{"id":"kyutai-stt","name":"Kyutai STT (Delayed Streams Modeling)","vendor":"Kyutai","category":"open-model","summary":"Open streaming STT models built for realtime: a 1B English/French model with 0.5 s delay and semantic VAD, and a 2.6B English model with 2.5 s delay. Served by a Rust server, PyTorch or MLX.","status":"GA","models":[{"name":"kyutai/stt-1b-en_fr","status":"Released","notes":"English + French, ~1B params, 0.5 s delay, semantic VAD."},{"name":"kyutai/stt-2.6b-en","status":"Released","notes":"English only, ~2.6B params, 2.5 s delay, higher accuracy."}],"transports":["WebSocket (moshi-server)","Python (PyTorch, MLX)"],"audio":{"input":"Handled by the provided scripts (mic or file); server streams PCM over WebSocket.","output":"n/a"},"languages":"English and French (1B) or English (2.6B).","latency":"Model-defined delay: 0.5 s (1B) or 2.5 s (2.6B) per README.","features":["word-level timestamps","semantic VAD (1B)","batched serving (README: an H100 handles hundreds of streams)","Apple Silicon via MLX"],"pricing":{"model":"free","items":[{"what":"Open weights","price":"$0","unit":"","notes":"Compute only."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Self-host compute only."},"free_tier":"Open weights.","source":"https://github.com/kyutai-labs/delayed-streams-modeling"},"limits":["Fixed algorithmic delay per model; 2.6B is too slow-to-final for snappy voice agents."],"regions":"Self-hosted.","setup":{"steps":["Quick test on a file: uvx --with moshi python -m moshi.run_inference --hf-repo kyutai/stt-2.6b-en audio.mp3","Production: cargo install --features cuda moshi-server, then moshi-server worker --config configs/config-stt-en_fr-hf.toml","Stream the mic to the server with scripts/stt_from_mic_rust_server.py (or stt_from_mic_mlx.py on a Mac)."],"endpoint":"ws://localhost:8080 (moshi-server default; check config)","auth":"Configurable API key in moshi-server config (check repo).","snippet_lang":"bash","snippet":"git clone https://github.com/kyutai-labs/delayed-streams-modeling && cd delayed-streams-modeling\ncargo install --features cuda moshi-server\nmoshi-server worker --config configs/config-stt-en_fr-hf.toml &\nuv run scripts/stt_from_mic_rust_server.py   # prints words live from your microphone\n# Mac without a server: uv run scripts/stt_from_mic_mlx.py"},"warnings":[{"severity":"medium","title":"Model weights are CC-BY-4.0","detail":"Code is Apache-2.0 but the model weights on Hugging Face are CC-BY-4.0, which requires attribution in your product."},{"severity":"medium","title":"Only English and French","detail":"No other languages are supported by the released STT models."},{"severity":"medium","title":"Repo activity slowed","detail":"The GitHub repo's last push was January 2026; expect community support rather than a vendor roadmap."},{"severity":"low","title":"Rust server build","detail":"The production server is a Rust crate built with CUDA features; plan build time and CUDA toolchain on the host."}],"best_for":"Self-hosted English/French realtime captions or voice agents that want built-in semantic VAD.","open_source":true,"self_hostable":true,"hardware":"NVIDIA GPU for moshi-server (README cites H100 and L40S batch sizes); Apple Silicon via MLX for local use.","license":"Code Apache-2.0; model weights CC-BY-4.0.","compliance":"Self-hosted.","docs":[{"label":"GitHub","url":"https://github.com/kyutai-labs/delayed-streams-modeling"},{"label":"stt-1b-en_fr","url":"https://huggingface.co/kyutai/stt-1b-en_fr"},{"label":"stt-2.6b-en","url":"https://huggingface.co/kyutai/stt-2.6b-en"}],"sources":["https://raw.githubusercontent.com/kyutai-labs/delayed-streams-modeling/main/README.md","https://api.github.com/repos/kyutai-labs/delayed-streams-modeling","https://huggingface.co/api/models/kyutai/stt-1b-en_fr","https://huggingface.co/api/models/kyutai/stt-2.6b-en"],"confidence":"medium","unverified":"Default server port/auth and exact streams-per-GPU figure (README text truncated in fetch).","cat":"open","kind":"stt","verified_at":"2026-10-10","facts":{"latency_ms":500,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":1,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"CC-BY-4.0","full_duplex":null,"streaming":true,"_notes":"Values for stt-1b-en_fr (0.5 s delay); stt-2.6b-en is English only with 2.5 s delay. Code Apache-2.0, weights CC-BY-4.0. Apple Silicon via MLX."}},{"id":"moonshine","name":"Moonshine Voice (Moonshine v2 streaming)","vendor":"Moonshine AI (Useful Sensors)","category":"open-model","summary":"MIT-licensed on-device STT library with streaming models (Moonshine v2, 2026) that do work while the user is still talking. Runs on CPU, phones, browsers (WASM) and Raspberry Pi.","status":"GA","models":[{"name":"Moonshine v2 streaming models","status":"Released 2026","notes":"Sliding-window attention encoder for bounded latency (arXiv 2602.12241). Sizes range from tiny (~1 MB micro) up to models the README says beat Whisper Large V3."},{"name":"Legacy non-streaming non-English models","status":"Legacy","notes":"Under a non-commercial Moonshine licence (exception to MIT)."}],"transports":["Local library (Python, JS/WASM, iOS, Android, desktop)"],"audio":{"input":"Microphone or PCM via the library.","output":"n/a"},"languages":"English plus additional languages (list on docs site; not captured).","latency":"Paper: bounded time-to-first-token independent of utterance length; third-party claim of sub-200 ms on edge devices is unverified.","features":["streaming partials","on-device","intent recognition and TTS in the same library","cross-platform"],"pricing":{"model":"free","items":[{"what":"Library and models","price":"$0","unit":"","notes":"MIT, except legacy non-English non-streaming models."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Runs on the user's device."},"free_tier":"Free and open source.","source":"https://github.com/moonshine-ai/moonshine"},"limits":["Accuracy depends on chosen model size and device CPU."],"regions":"On-device.","setup":{"steps":["pip install moonshine-voice","moonshine-voice mic --language en (live mic test)","See moonshine-voice.readthedocs.io for Python/JS/mobile APIs."],"endpoint":"n/a (local)","auth":"none","snippet_lang":"bash","snippet":"pip install moonshine-voice\nmoonshine-voice mic --language en   # live transcription from the default microphone"},"warnings":[{"severity":"medium","title":"Licence exception","detail":"Everything is MIT except the legacy non-streaming models for non-English languages, which stay under a non-commercial Moonshine licence. Use the streaming models for commercial products."},{"severity":"medium","title":"Benchmarks disagree","detail":"The arXiv abstract and the newer PDF on moonshine.ai give different accuracy/speed claims (6x vs 4x model size parity). Test on your own audio and hardware."},{"severity":"low","title":"Small models trade accuracy","detail":"Tiny and micro models suit commands and captions on weak hardware; long-form dictation needs the larger models."}],"best_for":"Private, offline, low-latency captions and voice commands on phones, browsers and single-board computers.","open_source":true,"self_hostable":true,"hardware":"CPU-only friendly: laptops, phones, Raspberry Pi, browser WASM.","license":"MIT (code and streaming models); legacy non-English non-streaming models non-commercial.","compliance":"On-device processing.","docs":[{"label":"GitHub","url":"https://github.com/moonshine-ai/moonshine"},{"label":"Docs","url":"https://moonshine-voice.readthedocs.io/"},{"label":"Moonshine v2 paper","url":"https://arxiv.org/abs/2602.12241"}],"sources":["https://raw.githubusercontent.com/moonshine-ai/moonshine/main/README.md","https://api.github.com/repos/moonshine-ai/moonshine","https://arxiv.org/abs/2602.12241","https://download.moonshine.ai/docs/moonshine_streaming_paper.pdf"],"confidence":"medium","unverified":"Supported language list and per-device latency numbers.","cat":"open","kind":"stt","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":true,"commercial_ok":true,"license_short":"MIT","full_duplex":null,"streaming":true,"_notes":"On-device library; model sizes range from tiny to large. English plus other languages (list not captured). Legacy non-English non-streaming models are non-commercial."}},{"id":"whisperlive","name":"WhisperLive (Collabora)","vendor":"Collabora (open source)","category":"open-model","summary":"Self-hosted near-live Whisper transcription server with WebSocket clients, VAD, and faster-whisper, TensorRT or OpenVINO backends.","status":"GA","models":[{"name":"Any Whisper size via faster_whisper / tensorrt / openvino backends","status":"Active project","notes":"Last push 2026-10-07."}],"transports":["WebSocket","REST (optional)"],"audio":{"input":"Client captures mic or file and streams 16 kHz float audio.","output":"n/a"},"languages":"Whisper's ~99 languages (model-dependent).","latency":"Depends on model size and GPU; Whisper is not natively streaming, so partials are re-decoded windows.","features":["VAD","multiple concurrent clients (max_clients)","max_connection_time","browser extension and iOS clients in repo","translation"],"pricing":{"model":"free","items":[{"what":"Software","price":"$0","unit":"","notes":"Compute only."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Self-host compute only."},"free_tier":"Open source.","source":"https://github.com/collabora/WhisperLive"},"limits":["Server-side max_clients and max_connection_time settings cap concurrency and session length."],"regions":"Self-hosted.","setup":{"steps":["pip install whisper-live","python3 run_server.py --port 9090 --backend faster_whisper","Connect with the Python client (below) or the browser extension."],"endpoint":"ws://localhost:9090","auth":"None by default; put it behind your own auth/TLS proxy.","snippet_lang":"python","snippet":"# server: python3 run_server.py --port 9090 --backend faster_whisper\nfrom whisper_live.client import TranscriptionClient  # pip install whisper-live\n\nclient = TranscriptionClient(\"localhost\", 9090, lang=\"en\", model=\"small\", use_vad=True)\nclient()  # streams the default microphone and prints segments"},"warnings":[{"severity":"medium","title":"No auth out of the box","detail":"The WebSocket server has no authentication; exposing port 9090 publicly lets anyone use your GPU. Put it behind a TLS proxy with auth."},{"severity":"medium","title":"Whisper is not a true streaming model","detail":"Partials come from repeatedly decoding a sliding window, which costs more GPU per stream and can rewrite earlier words."},{"severity":"low","title":"TensorRT backend is fiddly","detail":"The project recommends the Docker setup for TensorRT; native builds are version-sensitive."}],"best_for":"Self-hosted multilingual live captions on your own GPU with Whisper accuracy.","open_source":true,"self_hostable":true,"hardware":"NVIDIA GPU recommended for small and larger models; CPU possible with tiny/base or OpenVINO.","license":"MIT","compliance":"Self-hosted.","docs":[{"label":"GitHub","url":"https://github.com/collabora/WhisperLive"}],"sources":["https://raw.githubusercontent.com/collabora/WhisperLive/main/README.md","https://api.github.com/repos/collabora/WhisperLive"],"confidence":"medium","unverified":"Client constructor arguments taken from the README pattern; not executed.","cat":"open","kind":"stt","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":99,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":true,"commercial_ok":true,"license_short":"MIT","full_duplex":null,"streaming":true,"_notes":"Server around any Whisper size; partials are re-decoded windows, not native streaming. CPU workable with tiny/base or OpenVINO. Session length and client count are server settings."}},{"id":"whisperlivekit","name":"WhisperLiveKit","vendor":"Open source (QuentinFuxa)","category":"open-model","summary":"Popular self-hosted realtime transcription server with a web UI, combining SimulStreaming / WhisperStreaming policies with optional streaming speaker diarization (Sortformer).","status":"GA","models":[{"name":"Whisper models (tiny..large-v3) with SimulStreaming or LocalAgreement policies","status":"Active project","notes":"Last push 2026-10-01; CLI 'wlk'."}],"transports":["WebSocket","HTTP (web UI)"],"audio":{"input":"Browser mic via built-in web page or WebSocket clients.","output":"n/a"},"languages":"Whisper languages.","latency":"Policy and model dependent.","features":["streaming partials","realtime speaker diarization (Streaming Sortformer option)","web UI","model management CLI (wlk pull/rm/models)","file transcription to SRT"],"pricing":{"model":"free","items":[{"what":"Software","price":"$0","unit":"","notes":"Compute only."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Self-host compute only."},"free_tier":"Open source.","source":"https://github.com/QuentinFuxa/WhisperLiveKit"},"limits":["Single-server design; scale-out is up to you."],"regions":"Self-hosted.","setup":{"steps":["pip install whisperlivekit","wlk --model base --language en","Open the local web UI and allow microphone access."],"endpoint":"http://localhost:8000 (default UI; check CLI output)","auth":"None by default.","snippet_lang":"bash","snippet":"pip install whisperlivekit\nwlk --model base --language en      # starts server + web UI\nwlk transcribe --format srt podcast.mp3 -o podcast.srt   # offline file mode"},"warnings":[{"severity":"medium","title":"No built-in auth","detail":"Run it behind a reverse proxy with authentication before exposing it beyond localhost."},{"severity":"medium","title":"Diarization adds GPU load","detail":"Streaming diarization models run alongside Whisper; size the GPU accordingly."},{"severity":"low","title":"Default port unverified","detail":"Check the CLI output for the actual port before wiring clients."}],"best_for":"Quick self-hosted live transcription with a usable web UI and optional diarization.","open_source":true,"self_hostable":true,"hardware":"GPU recommended beyond the base model; Apple Silicon supported via MLX backends (per project, not verified).","license":"Apache-2.0","compliance":"Self-hosted.","docs":[{"label":"GitHub","url":"https://github.com/QuentinFuxa/WhisperLiveKit"}],"sources":["https://raw.githubusercontent.com/QuentinFuxa/WhisperLiveKit/main/README.md","https://api.github.com/repos/QuentinFuxa/WhisperLiveKit"],"confidence":"medium","unverified":"Default port, MLX support.","cat":"open","kind":"stt","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":99,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":true,"_notes":"Whisper models with SimulStreaming or LocalAgreement policies; languages follow Whisper (about 99). Optional streaming diarization."}},{"id":"whisper-streaming-simulstreaming","name":"whisper_streaming and SimulStreaming (UFAL)","vendor":"Charles University UFAL (open source)","category":"open-model","summary":"Research implementations that turn Whisper into a streaming transcriber. whisper_streaming (LocalAgreement policy, 2023) is now marked outdated by its authors in favour of SimulStreaming.","status":"GA","models":[{"name":"whisper_streaming","status":"Superseded","notes":"README: 'becoming outdated, replaced by SimulStreaming'. Last push 2025-11."},{"name":"SimulStreaming","status":"Active","notes":"Last push 2026-07; simultaneous policy for Whisper."}],"transports":["TCP/stdin server scripts (research code)"],"audio":{"input":"16 kHz mono audio via the provided server/clients.","output":"n/a"},"languages":"Whisper languages.","latency":"Configurable chunk/min-chunk; papers report a few seconds of latency on GPU (not re-verified).","features":["streaming transcription and translation","LocalAgreement or simultaneous decoding policies","VAD option"],"pricing":{"model":"free","items":[{"what":"Software","price":"$0","unit":"","notes":"Compute only."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Self-host compute only."},"free_tier":"Open source.","source":"https://github.com/ufal/SimulStreaming"},"limits":["Research-grade code: one stream per process, minimal ops tooling."],"regions":"Self-hosted.","setup":{"steps":["git clone https://github.com/ufal/SimulStreaming","Install requirements and follow README to run the server and stream audio.","Prefer wrappers such as WhisperLiveKit for a ready server."],"endpoint":"Local TCP server (see README)","auth":"none","snippet_lang":"bash","snippet":""},"warnings":[{"severity":"medium","title":"Original repo is outdated","detail":"The whisper_streaming authors say it is being replaced by SimulStreaming; start new work on SimulStreaming or a wrapper that bundles it."},{"severity":"medium","title":"Not production packaging","detail":"These are research implementations without auth, multi-tenant scaling or monitoring."},{"severity":"low","title":"Small community for SimulStreaming","detail":"SimulStreaming has far fewer stars and contributors than WhisperLive or WhisperLiveKit."}],"best_for":"Researchers and builders who want the underlying streaming policies to embed in their own server.","open_source":true,"self_hostable":true,"hardware":"GPU recommended for large Whisper models.","license":"MIT (both repos per GitHub metadata).","compliance":"Self-hosted.","docs":[{"label":"whisper_streaming","url":"https://github.com/ufal/whisper_streaming"},{"label":"SimulStreaming","url":"https://github.com/ufal/SimulStreaming"}],"sources":["https://raw.githubusercontent.com/ufal/whisper_streaming/main/README.md","https://api.github.com/repos/ufal/whisper_streaming","https://api.github.com/repos/ufal/SimulStreaming"],"confidence":"medium","unverified":"Exact run commands and latency figures (no snippet given).","cat":"open","kind":"stt","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":99,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":false,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"MIT","full_duplex":null,"streaming":true,"_notes":"Research code over raw TCP/stdin; papers report a few seconds of latency. Languages follow Whisper (about 99). whisper_streaming is superseded by SimulStreaming."}},{"id":"faster-whisper","name":"faster-whisper (engine for DIY streaming)","vendor":"SYSTRAN (open source)","category":"open-model","summary":"CTranslate2 re-implementation of Whisper, 4x-class faster than openai-whisper. Not a streaming server itself; it is the decoding engine most self-hosted live tools (WhisperLive, WhisperLiveKit) build on.","status":"GA","models":[{"name":"Whisper tiny..large-v3, large-v3-turbo, distil variants (CTranslate2 conversions)","status":"Active project","notes":"Last push 2026-10-10."}],"transports":["Python library"],"audio":{"input":"Files or numpy arrays (16 kHz).","output":"n/a"},"languages":"Whisper languages.","latency":"Batch engine; latency depends on chunking strategy you build.","features":["built-in Silero VAD filter","word timestamps","int8/float16 quantization","batched inference pipeline"],"pricing":{"model":"free","items":[{"what":"Software","price":"$0","unit":"","notes":"Compute only."}],"est_per_minute_usd":{"low":0,"high":0,"basis":"Self-host compute only."},"free_tier":"Open source.","source":"https://github.com/SYSTRAN/faster-whisper"},"limits":["No streaming API: you must buffer audio and re-transcribe windows (or use a wrapper)."],"regions":"Self-hosted.","setup":{"steps":["pip install faster-whisper","Load a model on CUDA or CPU (int8).","Buffer live audio into short windows, transcribe each, and merge (or use WhisperLive / WhisperLiveKit)."],"endpoint":"n/a (library)","auth":"none","snippet_lang":"python","snippet":"from faster_whisper import WhisperModel  # pip install faster-whisper\n\nmodel = WhisperModel(\"large-v3-turbo\", device=\"cuda\", compute_type=\"float16\")\n# Transcribe one buffered window of live audio (e.g. the last few seconds)\nsegments, info = model.transcribe(\"window.wav\", vad_filter=True, beam_size=1)\nfor s in segments:\n    print(f\"[{s.start:.1f}-{s.end:.1f}] {s.text}\")"},"warnings":[{"severity":"medium","title":"Not streaming by itself","detail":"faster-whisper only transcribes buffers; partial results, stabilization and endpointing are your job (or a wrapper's)."},{"severity":"medium","title":"Hallucination on silence","detail":"Whisper can invent text on silent or noisy windows; keep vad_filter on and drop low-probability segments."},{"severity":"low","title":"CUDA/cuDNN version coupling","detail":"GPU use depends on matching CTranslate2, CUDA and cuDNN versions; mismatches fail at load time."}],"best_for":"Building your own Whisper-based live pipeline or powering WhisperLive/WhisperLiveKit.","open_source":true,"self_hostable":true,"hardware":"NVIDIA GPU for large models in real time; CPU int8 workable for tiny/base/small.","license":"MIT (code); Whisper weights MIT.","compliance":"Self-hosted.","docs":[{"label":"GitHub","url":"https://github.com/SYSTRAN/faster-whisper"}],"sources":["https://api.github.com/repos/SYSTRAN/faster-whisper"],"confidence":"medium","unverified":"large-v3-turbo model alias availability in the installed version.","cat":"open","kind":"stt","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":99,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":true,"commercial_ok":true,"license_short":"MIT","full_duplex":null,"streaming":false,"_notes":"Batch decoding library, no streaming API by itself. CPU int8 workable for tiny/base/small. Languages follow Whisper (about 99)."}},{"id":"elevenlabs","name":"ElevenLabs TTS API","vendor":"ElevenLabs","category":"tts","summary":"The best-known voice API: Flash v2.5 for low-latency agents over a text-input WebSocket, plus the new Eleven v4 / v4 Turbo (launched 2026-09-28) that stream through a separate Text to Dialogue WebSocket.","status":"GA","models":[{"name":"eleven_v4","status":"GA (flagship, launched 2026-09-28)","notes":"Highest quality, 90+ languages, 10,000 chars/request. Used via the Text to Dialogue API, not the classic TTS WebSocket."},{"name":"eleven_v4_turbo","status":"GA (launched 2026-09-28)","notes":"~100 ms median inference (vendor). For agents via wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input. One voice per connection."},{"name":"eleven_flash_v2_5","status":"GA","notes":"~75 ms model latency (vendor), 32 languages, 40,000 chars/request. Recommended for the classic stream-input WebSocket."},{"name":"eleven_flash_v2","status":"GA","notes":"English only, ~75 ms (vendor)."},{"name":"eleven_v3 / eleven_v3_conversational","status":"Previous generation","notes":"70+ languages; v3 conversational ~280 ms. Not accepted on the classic TTS WebSocket; use Text to Dialogue."},{"name":"eleven_multilingual_v2","status":"Previous generation","notes":"29 languages, 10,000 chars/request. Default model_id of the stream-input WebSocket if you omit it."},{"name":"eleven_turbo_v2_5 / eleven_turbo_v2","status":"Deprecated","notes":"Docs say use Flash instead; functionally equivalent."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text (SSML parsing optional via enable_ssml_parsing on the WebSocket)","output":"MP3 by default; output_format values follow codec_samplerate_bitrate, e.g. mp3_44100_128, pcm_16000/22050/24000/44100, ulaw_8000, alaw_8000, opus_48000_* (format list from third-party mirrors of the API reference; some higher-quality formats are tier-gated)"},"languages":"Flash v2.5: 32; v3: 70+; v4: 90+","voices":"Large shared voice library; instant and professional voice cloning: yes","latency":"Vendor claims: Flash v2.5 ~75 ms model latency, v4 Turbo ~100 ms median inference, v3 conversational ~280 ms. All exclude network and application latency.","features":["input streaming (WebSocket)","multi-context WebSocket for barge-in","character alignment / timestamps (alignment, sync_alignment)","voice cloning","voice design","pronunciation dictionaries","data residency endpoints (US, EU, India, Singapore)"],"pricing":{"model":"per-character","items":[{"what":"Flash / Turbo","price":"$0.04","unit":"per 1K characters","notes":"API pricing page; ~$0.04/min per vendor"},{"what":"Eleven v4 Turbo","price":"$0.011","unit":"per 1K characters","notes":"Promotional, discounted from $0.04 until Oct 12 (2026)"},{"what":"Eleven v4","price":"$0.022","unit":"per 1K characters","notes":"Promotional, discounted from $0.08 until Oct 12 (2026)"},{"what":"Eleven v3","price":"$0.08","unit":"per 1K characters","notes":""},{"what":"Eleven v3 Conversational","price":"$0.04","unit":"per 1K characters","notes":""},{"what":"Multilingual v2","price":"$0.08","unit":"per 1K characters","notes":""},{"what":"Starter plan","price":"$6/month ($1 first month)","unit":"monthly","notes":"150,000 Flash chars or 75,000 v3 chars"},{"what":"Creator plan","price":"$22/month","unit":"monthly","notes":"550,000 Flash chars"},{"what":"Pro plan","price":"$99/month","unit":"monthly","notes":"2,475,000 Flash chars"},{"what":"Scale plan","price":"$299/month","unit":"monthly","notes":"7,475,000 Flash chars"},{"what":"Business plan","price":"$990/month","unit":"monthly","notes":"24,750,000 Flash chars"}],"est_per_minute_usd":{"low":0.0099,"high":0.072,"basis":"900 chars/min. Low = v4 Turbo promo rate ($0.011/1K, ends Oct 12 2026; $0.036/min after). Flash v2.5 = $0.036/min. High = v3 or Multilingual v2 at $0.08/1K."},"free_tier":"Free / pay-as-you-go: 10,000-20,000 characters depending on model. Commercial-use terms on the free tier not re-verified; check the plan terms.","source":"https://elevenlabs.io/pricing/api"},"limits":["Chars per request: Flash v2.5 40,000; Multilingual v2 and v4 10,000; v3 5,000 (models page)","WebSocket inactivity_timeout default 20 s, max 180 s","Third-party reports (Vapi support): max 5 simultaneous contexts per multi-context WebSocket; not confirmed on official pages","Plan concurrency limits are not shown on the API pricing page; third-party lists (Free 2 ... Business 15) are unofficial","Text to Dialogue socket waits for ~40 characters and 8 words before emitting audio unless you flush"],"regions":"Global plus residency hosts: api.us.elevenlabs.io, api.eu.residency.elevenlabs.io, api.in.residency.elevenlabs.io, api.sg.residency.elevenlabs.io","setup":{"steps":["Create an API key in the ElevenLabs dashboard.","Pick a voice_id from the voice library.","For LLM token streaming with Flash: open wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input?model_id=eleven_flash_v2_5 with the xi-api-key header.","Send {\"text\":\" \"} first, then text chunks ending in a space, then {\"text\":\"\"} to finish.","For v4 Turbo use the Text to Dialogue WebSocket instead (different message format: register voices in the first message)."],"endpoint":"wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input","auth":"xi-api-key header (or authorization bearer / single_use_token query param for clients)","snippet_lang":"javascript","snippet":"// npm i ws  - streams LLM-style text chunks in, saves raw PCM out\nimport WebSocket from \"ws\";\nimport fs from \"fs\";\n\nconst voiceId = \"YOUR_VOICE_ID\";\nconst url = `wss://api.elevenlabs.io/v1/text-to-speech/${voiceId}/stream-input?model_id=eleven_flash_v2_5&output_format=pcm_24000`;\nconst ws = new WebSocket(url, { headers: { \"xi-api-key\": process.env.ELEVENLABS_API_KEY } });\nconst out = fs.createWriteStream(\"out_24k_s16le.pcm\");\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({ text: \" \" }));                 // init message\n  for (const t of [\"Hello there. \", \"This text arrives \", \"in pieces, like LLM tokens. \"]) {\n    ws.send(JSON.stringify({ text: t }));\n  }\n  ws.send(JSON.stringify({ text: \"\" }));                  // end of input\n});\nws.on(\"message\", (raw) => {\n  const msg = JSON.parse(raw.toString());\n  if (msg.audio) out.write(Buffer.from(msg.audio, \"base64\"));\n  if (msg.isFinal) ws.close();\n});\nws.on(\"close\", () => out.end());"},"warnings":[{"severity":"high","title":"v3 and v4 do not work on the classic TTS WebSocket","detail":"The /text-to-speech/{voice_id}/stream-input socket rejects eleven_v3 and eleven_v4 models. Use eleven_flash_v2_5 there, or switch to the Text to Dialogue WebSocket (wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input) for eleven_v4_turbo, which has a different message format."},{"severity":"high","title":"v4 launch prices are promotional","detail":"The API pricing page shows v4 at $0.022/1K and v4 Turbo at $0.011/1K only until Oct 12 (2026); list prices are $0.08 and $0.04/1K. Budget on the list price."},{"severity":"medium","title":"Default model on the socket is not Flash","detail":"If you omit model_id the stream-input socket uses eleven_multilingual_v2 (higher latency, double the price of Flash). Always set model_id explicitly."},{"severity":"medium","title":"Buffering adds latency with small chunks","detail":"The server buffers text using chunk_length_schedule (default [120,160,250,290] chars). For conversational agents send flush:true at the end of each turn or tune the schedule, otherwise the first audio waits for 120 characters."},{"severity":"medium","title":"Turbo models are deprecated","detail":"eleven_turbo_v2_5 and eleven_turbo_v2 are marked deprecated; migrate to Flash."},{"severity":"low","title":"Request logging is on by default","detail":"enable_logging defaults to true. Zero-retention mode (enable_logging=false) is an enterprise feature per earlier docs; confirm before sending sensitive text."}],"best_for":"General-purpose voice agents and content with the largest voice library; Flash v2.5 for latency, v4 for quality.","open_source":false,"self_hostable":false,"compliance":"Data-residency endpoints for EU, India and Singapore exist. Certifications not re-verified in this pass.","docs":[{"label":"Models","url":"https://elevenlabs.io/docs/overview/models"},{"label":"TTS WebSocket reference","url":"https://elevenlabs.io/docs/api-reference/text-to-speech/v-1-text-to-speech-voice-id-stream-input"},{"label":"TTS vs Text to Dialogue WebSockets","url":"https://elevenlabs.io/docs/eleven-api/guides/how-to/websockets/tts-vs-ttd-websockets"},{"label":"v4 changelog","url":"https://elevenlabs.io/docs/changelog/2026/9/28"}],"sources":["https://elevenlabs.io/pricing/api","https://elevenlabs.io/docs/overview/models","https://elevenlabs.io/docs/changelog/2026/9/28","https://elevenlabs.io/docs/api-reference/text-to-speech/v-1-text-to-speech-voice-id-stream-input","https://elevenlabs.io/docs/eleven-api/guides/how-to/websockets/tts-vs-ttd-websockets"],"confidence":"high","unverified":"Per-plan concurrency limits; exact output_format list (taken from third-party mirrors); 5-contexts-per-socket limit (third-party); free-tier commercial terms.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"ElevenLabs","facts":{"cat":"tts","latency_ms":75,"languages":32,"max_session_min":null,"concurrency":6,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":40,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":true,"emotion_control":null,"ssml":true,"telephony_8k":true,"commercial_free_tier":null,"_notes":"Figures are for Flash v2.5, the model for the streaming WebSocket (~75 ms model latency, 32 languages, $0.04/1K). v4 covers 90+ languages ($0.08/1K list, promo $0.022 until Oct 12 2026). Concurrency 6 is Flash on the Starter plan (Pro 20). Timestamps are character-level alignment. SSML parsing optional on the socket.","_added":["concurrency: https://elevenlabs.io/docs/models"]}},{"id":"cartesia-sonic","name":"Cartesia Sonic","vendor":"Cartesia","category":"tts","summary":"State-space-model TTS built for voice agents; one WebSocket carries many contexts with continuation, word and phoneme timestamps, and native mu-law/a-law output. Current model is Sonic 3.6.","status":"GA","models":[{"name":"sonic-3.6","status":"GA (alias to latest stable snapshot)","notes":"44 languages; backwards compatible with Sonic 3.5."},{"name":"sonic-3.6-2026-08-27","status":"Stable snapshot","notes":"Pin this in production so behaviour does not change under you."},{"name":"sonic-3.5, sonic-3","status":"Older, still accepted on the WebSocket","notes":""},{"name":"sonic-latest / sonic-preview","status":"Beta","notes":"sonic-preview may change without notice; not for production."},{"name":"sonic-2, sonic-turbo, sonic","status":"Older models","notes":"Listed only as older models; not in the current WebSocket model_id list."}],"transports":["WebSocket","HTTP chunked","SSE"],"audio":{"input":"Text (no SSML required; break tags supported)","output":"Raw container; pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw at 8000/16000/22050/24000/44100/48000 Hz on the WebSocket"},"languages":"44 (Sonic 3.6)","voices":"Voice library plus instant voice cloning (Pro and up) and professional cloning (Startup and up)","latency":"Third-party sources quote Cartesia's claims of ~90 ms time-to-first-audio for Sonic 3 and ~40 ms for Sonic Turbo; Sonic 3.6 described as sub-90 ms. Independent benchmarks reported 128-166 ms medians including network for Sonic 3/3.5.","features":["input streaming via contexts (continue flag)","word timestamps","phoneme timestamps","emotion, speed and volume controls (generation_config)","voice cloning","pronunciation dictionaries","cancel per context","access tokens for browser clients"],"pricing":{"model":"subscription","items":[{"what":"Free","price":"$0","unit":"monthly","notes":"20K credits; TTS concurrency 2; no cloning"},{"what":"Pro","price":"$5/month","unit":"monthly","notes":"100K credits; TTS concurrency 3; instant cloning; overage $65 per 1M credits"},{"what":"Startup","price":"$49/month","unit":"monthly","notes":"1.25M credits; TTS concurrency 5; overage $45 per 1M credits"},{"what":"Scale","price":"$299/month","unit":"monthly","notes":"8M credits; TTS concurrency 15; overage $38 per 1M credits"},{"what":"TTS credit rate","price":"1 credit","unit":"per character","notes":"~750-800 credits per minute of audio (vendor FAQ); each break tag = 1 credit"}],"est_per_minute_usd":{"low":0.034,"high":0.059,"basis":"900 chars/min at 1 credit/char. Low = Scale plan effective ~$37.4/1M; high = Pro overage $65/1M."},"free_tier":"20K credits/month, overages blocked; commercial use not listed for Free (Pro and up: yes).","source":"https://cartesia.ai/pricing"},"limits":["TTS concurrency: Free 2, Pro 3, Startup 5, Scale 15","max_buffer_delay_ms default 3000 (0-5000)","voice and output_format must stay constant within a context","Overages must be enabled; Free stops at the credit cap"],"regions":"Not specified on the pages checked","setup":{"steps":["Create an API key in the Cartesia playground/console.","Pick a voice ID.","Open wss://api.cartesia.ai/tts/websocket with X-API-Key and a cartesia_version.","Send generation requests sharing one context_id; set continue:true on every chunk except the last.","Read base64 'chunk' messages until 'done'."],"endpoint":"wss://api.cartesia.ai/tts/websocket","auth":"X-API-Key header (server) or access_token query param (browser); cartesia_version required (example 2026-08-14)","snippet_lang":"python","snippet":"# pip install websockets   (streams text chunks into one context)\nimport asyncio, base64, json, os, websockets\n\nURL = \"wss://api.cartesia.ai/tts/websocket?cartesia_version=2026-08-14\"\nHDR = {\"X-API-Key\": os.environ[\"CARTESIA_API_KEY\"]}\n\nasync def main():\n    async with websockets.connect(URL, additional_headers=HDR) as ws:\n        chunks = [\"Hello there. \", \"This arrives \", \"token by token.\"]\n        for i, text in enumerate(chunks):\n            await ws.send(json.dumps({\n                \"model_id\": \"sonic-3.6\",\n                \"transcript\": text,\n                \"voice\": {\"mode\": \"id\", \"id\": \"YOUR_VOICE_ID\"},\n                \"output_format\": {\"container\": \"raw\", \"encoding\": \"pcm_s16le\", \"sample_rate\": 24000},\n                \"context_id\": \"turn-1\",\n                \"continue\": i < len(chunks) - 1,\n            }))\n        with open(\"out_24k_s16le.pcm\", \"wb\") as f:\n            async for raw in ws:\n                msg = json.loads(raw)\n                if msg.get(\"type\") == \"chunk\":\n                    f.write(base64.b64decode(msg[\"data\"]))\n                elif msg.get(\"type\") in (\"done\", \"error\"):\n                    break\n\nasyncio.run(main())"},"warnings":[{"severity":"medium","title":"Pin a dated snapshot","detail":"sonic-3.6 is an alias that moves to the newest stable snapshot. Pin sonic-3.6-2026-08-27 (or the current dated ID) for production so voice behaviour does not change silently."},{"severity":"medium","title":"Low TTS concurrency on small plans","detail":"Free allows 2 and Pro 3 concurrent TTS generations; a handful of simultaneous calls will hit 429s. Size the plan to peak simultaneous speakers, not monthly volume."},{"severity":"medium","title":"Last chunk must set continue:false","detail":"If you never send a final request with continue:false (or flush), the server waits up to max_buffer_delay_ms (default 3 s) before speaking the tail of the turn."},{"severity":"low","title":"Raw container only on the WebSocket","detail":"The socket returns headerless raw audio; you must know the encoding and sample rate to play or wrap it as WAV."},{"severity":"low","title":"Break tags cost credits","detail":"Each break tag counts as 1 credit; markup-heavy text costs slightly more than the spoken characters."}],"best_for":"Latency-critical voice agents and telephony (native mu-law 8 kHz), with timestamps for interruption handling.","open_source":false,"self_hostable":false,"compliance":"Not verified in this pass (enterprise custom terms available).","docs":[{"label":"Models","url":"https://docs.cartesia.ai/build-with-cartesia/tts-models/latest"},{"label":"WebSocket API","url":"https://docs.cartesia.ai/api-reference/tts/websocket"},{"label":"Pricing","url":"https://cartesia.ai/pricing"}],"sources":["https://cartesia.ai/pricing","https://docs.cartesia.ai/build-with-cartesia/tts-models/latest","https://docs.cartesia.ai/api-reference/tts/websocket","https://humannessindex.vapi.ai/models/cartesia-sonic-3-5"],"confidence":"medium","unverified":"Latency numbers (from third-party write-ups of Cartesia claims); whether cartesia_version must be header or query (snippet uses query); regions; voice count.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"Cartesia Sonic","facts":{"cat":"tts","latency_ms":90,"languages":44,"max_session_min":null,"concurrency":3,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":50,"voices":null,"voice_cloning":true,"instant_cloning":true,"text_streaming_input":true,"word_timestamps":true,"emotion_control":true,"ssml":null,"telephony_8k":true,"commercial_free_tier":false,"_notes":"Price is the Pro plan effective rate ($5 for 100K credits, 1 credit per char); Pro overage $65/1M, Scale about $37/1M. Latency ~90 ms is a vendor claim quoted by third parties. Concurrency 3 on Pro (Free 2, Scale 15). Break tags supported but not full SSML. Free plan not listed for commercial use."}},{"id":"deepgram-aura-2","name":"Deepgram Aura-2","vendor":"Deepgram","category":"tts","summary":"Enterprise-focused TTS tuned for clear, business-style voice agents. WebSocket input streaming with Speak/Flush/Clear/Close messages; generous concurrency on pay-as-you-go.","status":"GA","models":[{"name":"aura-2-<voice>-<lang> (e.g. aura-2-thalia-en)","status":"GA","notes":"Model and voice are one ID. English, Spanish, plus Dutch, French, German, Italian, Japanese added Dec 2025."},{"name":"aura-<voice>-en (Aura-1)","status":"GA (older)","notes":"Half the price of Aura-2."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Plain text","output":"WebSocket: linear16, mulaw, alaw (container none); linear16 8/16/24/32/48 kHz, mulaw/alaw 8 or 16 kHz. REST adds mp3 (22.05 kHz), opus (48 kHz ogg), flac, aac."},"languages":"7 for Aura-2: en, es, nl, fr, de, it, ja (changelog Dec 2025 / Jan 2026)","voices":"40+ English voices per Deepgram marketing; ~90 total across languages per a third-party catalog. Voice cloning: not offered on the pages checked.","latency":"Vendor claim: sub-200 ms (Aura-2 product page). No figure on the streaming docs page.","features":["input streaming (Speak + Flush)","Clear message for barge-in","self-hosted deployment option","telephony encodings"],"pricing":{"model":"per-character","items":[{"what":"Aura-2","price":"$0.030","unit":"per 1K characters","notes":"Pay As You Go"},{"what":"Aura-2","price":"$0.027","unit":"per 1K characters","notes":"Growth plan"},{"what":"Aura-1","price":"$0.0150 / $0.0135","unit":"per 1K characters","notes":"PAYG / Growth"}],"est_per_minute_usd":{"low":0.0243,"high":0.027,"basis":"900 chars/min; Aura-2 Growth vs PAYG"},"free_tier":"$200 credit for new accounts","source":"https://deepgram.com/pricing"},"limits":["2,000 characters per Speak payload (413 above that)","2,400 characters per minute throughput per socket","20 Flush messages per 60 s","60-minute max socket lifetime","Voice and output settings fixed per connection","TTS concurrency: PAYG 45, Growth 60 (REST + WSS combined)"],"regions":"Hosted API plus self-hosted/on-prem option","setup":{"steps":["Create an API key in the Deepgram console.","Connect to wss://api.deepgram.com/v1/speak with model, encoding and sample_rate query params.","Send {type:'Speak', text} messages as tokens arrive, then {type:'Flush'} at end of turn.","Write binary frames as audio; JSON frames are control/metadata (Flushed, Warning)."],"endpoint":"wss://api.deepgram.com/v1/speak","auth":"Authorization: Token <API_KEY> header","snippet_lang":"javascript","snippet":"// npm i ws\nimport WebSocket from \"ws\";\nimport fs from \"fs\";\n\nconst url = \"wss://api.deepgram.com/v1/speak?model=aura-2-thalia-en&encoding=linear16&sample_rate=24000\";\nconst ws = new WebSocket(url, { headers: { Authorization: `Token ${process.env.DEEPGRAM_API_KEY}` } });\nconst out = fs.createWriteStream(\"out_24k_s16le.pcm\");\n\nws.on(\"open\", () => {\n  for (const t of [\"Hello there. \", \"This text is streamed \", \"sentence by sentence.\"]) {\n    ws.send(JSON.stringify({ type: \"Speak\", text: t }));\n  }\n  ws.send(JSON.stringify({ type: \"Flush\" }));\n});\nws.on(\"message\", (data, isBinary) => {\n  if (isBinary) return out.write(data);\n  const msg = JSON.parse(data.toString());\n  if (msg.type === \"Flushed\") ws.send(JSON.stringify({ type: \"Close\" }));\n});\nws.on(\"close\", () => out.end());"},"warnings":[{"severity":"high","title":"Flush rate limit","detail":"Only 20 Flush messages per 60 s per socket. Flushing after every LLM token or clause will trigger warnings; flush once per turn and let sentence punctuation drive synthesis."},{"severity":"medium","title":"Streaming formats are limited","detail":"The WebSocket only outputs linear16, mulaw and alaw. MP3/Opus are REST-only, so browsers need a PCM player or you transcode."},{"severity":"medium","title":"One voice per connection","detail":"Model/voice and encoding are fixed at connect time. Switching voice mid-call means a new socket; Deepgram recommends one socket per conversation."},{"severity":"medium","title":"Throughput cap per socket","detail":"2,400 characters per minute per socket is fine for one live speaker but too slow for bulk narration; use REST for long-form."},{"severity":"low","title":"WAV headers cause clicks in telephony","detail":"For REST telephony output set container=none; WAV headers mid-stream produce audible clicks."}],"best_for":"Enterprise voice agents already on Deepgram STT; high concurrency on pay-as-you-go; self-hosting needs.","open_source":false,"self_hostable":true,"compliance":"Self-hosted deployment available (Deepgram docs). Certifications not re-verified here.","docs":[{"label":"Streaming TTS","url":"https://developers.deepgram.com/docs/streaming-text-to-speech"},{"label":"Media output settings","url":"https://developers.deepgram.com/docs/tts-media-output-settings"},{"label":"TTS changelog","url":"https://developers.deepgram.com/changelog/text-to-speech-changelog"}],"sources":["https://deepgram.com/pricing","https://developers.deepgram.com/docs/streaming-text-to-speech","https://developers.deepgram.com/docs/tts-media-output-settings","https://deepgram.com/learn/aura-2-now-speaks-dutch-french-german-italian-japanese"],"confidence":"high","unverified":"Exact voice count; latency is vendor marketing only.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":200,"languages":7,"max_session_min":60,"concurrency":45,"free_tier":true,"free_credit_usd":200,"hipaa":true,"soc2":true,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_1m_chars":30,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":true,"commercial_free_tier":null,"_notes":"Sub-200 ms vendor claim. $30/1M PAYG, $27/1M Growth; Aura-1 is half price. Concurrency 45 on PAYG (REST and WSS combined). Socket max 60 minutes, 2,400 chars per minute per socket, 20 flushes per minute. HIPAA and SOC 2 are Deepgram vendor claims, not re-verified."}},{"id":"openai-tts","name":"OpenAI Text-to-Speech (gpt-4o-mini-tts)","vendor":"OpenAI","category":"tts","summary":"Steerable TTS via /v1/audio/speech: you pass natural-language 'instructions' for tone and style. Output streams over chunked HTTP; there is no text-input WebSocket for TTS (use the Realtime API for that).","status":"GA","models":[{"name":"gpt-4o-mini-tts (alias -> gpt-4o-mini-tts-2025-12-15)","status":"GA","notes":"Default snapshot; supports instructions and custom voices. 2,000 max input tokens."},{"name":"gpt-4o-mini-tts-2025-03-20","status":"Older snapshot","notes":"Some developers report it follows style instructions better; community reports say it is being deprecated (date not confirmed)."},{"name":"tts-1","status":"GA (legacy)","notes":"Lower latency, lower quality; 9 voices."},{"name":"tts-1-hd","status":"GA (legacy)","notes":"Higher quality; 9 voices."}],"transports":["HTTP chunked","SSE"],"audio":{"input":"Text plus optional free-text instructions","output":"mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless)"},"languages":"Follows Whisper language support; voices optimised for English","voices":"13 built-in (alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, cedar); custom voices for eligible customers with a consent recording","latency":"No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes.","features":["output streaming","style instructions (tone, accent, whispering, speed)","custom voices (eligible customers, consent required)","OpenAI-compatible endpoint copied by many vendors"],"pricing":{"model":"per-token","items":[{"what":"gpt-4o-mini-tts text input","price":"$0.60","unit":"per 1M text tokens","notes":""},{"what":"gpt-4o-mini-tts audio output","price":"$12.00","unit":"per 1M audio tokens","notes":""},{"what":"tts-1","price":"$15.00","unit":"per 1M characters","notes":""},{"what":"tts-1-hd","price":"$30.00","unit":"per 1M characters","notes":""}],"est_per_minute_usd":{"low":0.0135,"high":0.027,"basis":"900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now)."},"free_tier":"None specific to TTS","source":"https://developers.openai.com/api/docs/pricing"},"limits":["gpt-4o-mini-tts max 2,000 input tokens per request","Rate limits by tier: Build 2,000 RPM / 150K TPM; Launch 10,000 RPM / 2M TPM; Grow 10,000 RPM / 8M TPM"],"regions":"Not specified","setup":{"steps":["Create an API key.","Call POST https://api.openai.com/v1/audio/speech with model, voice, input and optional instructions.","Use response_format pcm or wav and iterate the streamed body.","Chunk long LLM output into sentences yourself (no input streaming)."],"endpoint":"https://api.openai.com/v1/audio/speech","auth":"Authorization: Bearer <OPENAI_API_KEY>","snippet_lang":"python","snippet":"# pip install openai\nfrom openai import OpenAI\n\nclient = OpenAI()  # reads OPENAI_API_KEY\n\nwith client.audio.speech.with_streaming_response.create(\n    model=\"gpt-4o-mini-tts\",\n    voice=\"marin\",\n    input=\"Hello! Your order has shipped and should arrive on Tuesday.\",\n    instructions=\"Warm, upbeat customer-support tone.\",\n    response_format=\"pcm\",  # 24 kHz 16-bit mono, no header\n) as resp:\n    with open(\"out_24k_s16le.pcm\", \"wb\") as f:\n        for chunk in resp.iter_bytes(4096):\n            f.write(chunk)  # or feed to a player as it arrives"},"warnings":[{"severity":"high","title":"No text-input streaming","detail":"/v1/audio/speech needs the full input up front. For LLM token streams you must split into sentences and fire one request per sentence (watch rate limits), or use the Realtime API instead."},{"severity":"medium","title":"Instruction-following varies by snapshot","detail":"Developers report the 2025-12-15 snapshot follows style instructions (e.g. whispering) less consistently than 2025-03-20. Pin a dated snapshot and test your prompts."},{"severity":"medium","title":"Token billing makes cost harder to predict","detail":"gpt-4o-mini-tts bills audio output tokens, so slower speech or long pauses cost more than the character count suggests. Measure on your own content."},{"severity":"medium","title":"Disclosure is required","detail":"OpenAI usage policies require telling end users the voice is AI-generated. Custom voices need a recorded consent statement from the speaker."},{"severity":"low","title":"Legacy models have fewer voices","detail":"tts-1 and tts-1-hd support only 9 voices; marin and cedar are gpt-4o-mini-tts only."}],"best_for":"Teams already on OpenAI that want steerable, prompt-styled speech for non-interactive or sentence-chunked use.","open_source":false,"self_hostable":false,"compliance":"OpenAI platform terms; disclosure of AI voice required by usage policy.","docs":[{"label":"TTS guide","url":"https://developers.openai.com/api/docs/guides/text-to-speech"},{"label":"Model page","url":"https://developers.openai.com/api/docs/models/gpt-4o-mini-tts"}],"sources":["https://developers.openai.com/api/docs/pricing","https://developers.openai.com/api/docs/guides/text-to-speech","https://developers.openai.com/api/docs/models/gpt-4o-mini-tts","https://developers.openai.com/docs/changelog"],"confidence":"high","unverified":"Per-minute cost for gpt-4o-mini-tts; deprecation date of the 2025-03-20 snapshot; SSE stream_format option.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"OpenAI TTS","facts":{"cat":"tts","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":15,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":false,"word_timestamps":null,"emotion_control":true,"ssml":null,"telephony_8k":false,"commercial_free_tier":null,"_notes":"$15/1M is tts-1 (tts-1-hd $30/1M). gpt-4o-mini-tts is token-billed ($12 per 1M audio tokens). HTTP chunked and SSE only; full text needed up front. tts-1 has 9 voices, gpt-4o-mini-tts has more. Custom voices limited to eligible customers with consent. No mulaw output (PCM is 24 kHz)."}},{"id":"google-cloud-tts","name":"Google Cloud Text-to-Speech (Chirp 3 HD and Gemini-TTS)","vendor":"Google Cloud","category":"tts","summary":"Cloud TTS offers Chirp 3 HD voices with bidirectional gRPC streaming (text in, audio out) and prompt-steerable Gemini-TTS models billed per audio token. Newest Gemini 3.8 Flash TTS models are Preview and only on the Gemini Enterprise API.","status":"GA","models":[{"name":"Chirp 3: HD (e.g. en-US-Chirp3-HD-Charon)","status":"GA","notes":"30 voices, 53 locales; the only voice type supported by StreamingSynthesize (bidi streaming, Preview feature)."},{"name":"gemini-2.5-flash-tts","status":"GA","notes":"Single and multi-speaker; streaming output PCM/ALAW/MULAW/OGG_OPUS."},{"name":"gemini-2.5-pro-tts","status":"GA","notes":"Higher control for podcasts/audiobooks."},{"name":"gemini-2.5-flash-lite-preview-tts","status":"Preview","notes":"Single speaker."},{"name":"gemini-3.1-flash-tts-preview","status":"Preview","notes":"Inline tags like [laughs], [sigh]."},{"name":"Gemini 3.8 Flash TTS / 3.8 Flash-Lite TTS","status":"Preview (Gemini Enterprise API only)","notes":"Not available through the Cloud TTS API; includes voice design and voice replication."},{"name":"Instant Custom Voice (Chirp 3)","status":"GA (restricted)","notes":"$60/1M characters."}],"transports":["gRPC","HTTP chunked"],"audio":{"input":"Text, SSML (legacy voices), natural-language prompt for Gemini-TTS","output":"Streaming: PCM (default), ALAW, MULAW, OGG_OPUS. Batch: LINEAR16, ALAW, MULAW, MP3, OGG_OPUS, PCM."},"languages":"Chirp 3 HD: 53 locales; Gemini-TTS: see per-model list","voices":"30 named Chirp 3 HD / Gemini voices (Achernar ... Zubenelgenubi); Instant Custom Voice cloning available","latency":"No numeric claim on the pages checked; Gemini-TTS described as 'very low latency'.","features":["bidirectional text-in/audio-out streaming (Chirp 3 HD, Preview)","prompt-based style control (Gemini-TTS)","multi-speaker dialogue (Gemini-TTS)","regional endpoints (global, us, eu, asia-southeast1, europe-west2, asia-northeast1 for Chirp 3 HD)","instant custom voice"],"pricing":{"model":"per-character","items":[{"what":"Chirp 3: HD voices","price":"$30","unit":"per 1M characters","notes":"First 1M characters/month free"},{"what":"Instant custom voice","price":"$60","unit":"per 1M characters","notes":"No free tier"},{"what":"Gemini 2.5 Flash TTS / 2.5 Flash-Lite Preview TTS","price":"$0.50 in / $10.00 out","unit":"per 1M text tokens / per 1M audio tokens","notes":"Audio = 25 tokens per second"},{"what":"Gemini 2.5 Pro TTS","price":"$1.00 in / $20.00 out","unit":"per 1M tokens","notes":""},{"what":"Gemini 3.1 Flash TTS (Preview)","price":"$1.00 in / $20.00 out","unit":"per 1M tokens","notes":""},{"what":"Gemini 3.8 Flash TTS (Preview)","price":"$0.50 in / $9.00 out","unit":"per 1M tokens","notes":"Through Dec 31 2026; $1.00 / $18.00 from Jan 1 2027"},{"what":"Gemini 3.8 Flash-Lite TTS (Preview)","price":"$0.50 in / $6.00 out","unit":"per 1M tokens","notes":"Through Dec 31 2026; $1.00 / $12.00 from Jan 1 2027"},{"what":"Neural2 / WaveNet / Standard / Studio (legacy)","price":"$16 / $4 / $4 / $160","unit":"per 1M characters","notes":"Not low-latency streaming voices"}],"est_per_minute_usd":{"low":0.009,"high":0.03,"basis":"Chirp 3 HD: 900 chars x $30/1M = $0.027. Gemini: 60 s x 25 tokens = 1,500 audio tokens/min; 3.8 Flash-Lite $0.009, 3.8 Flash $0.0135 (promo), 2.5 Flash $0.015, 2.5 Pro $0.03 (text input cost negligible)."},"free_tier":"Chirp 3 HD: 1M characters/month; WaveNet/Standard 4M; Gemini-TTS: none","source":"https://cloud.google.com/text-to-speech/pricing"},"limits":["Gemini-TTS: 8,192 input tokens, 16,384 output tokens per request","StreamingSynthesize: first message must be config only; Preview (Pre-GA terms)","Quotas per project; see quotas page"],"regions":"Chirp 3 HD GA in global, us, eu, asia-southeast1, europe-west2, asia-northeast1","setup":{"steps":["Enable the Cloud Text-to-Speech API in a GCP project with billing.","Authenticate with Application Default Credentials (gcloud auth application-default login or a service account).","pip install google-cloud-texttospeech.","Call streaming_synthesize with a config message followed by text messages."],"endpoint":"texttospeech.googleapis.com (gRPC StreamingSynthesize)","auth":"Google Cloud IAM / Application Default Credentials (OAuth), not a simple API key for streaming","snippet_lang":"python","snippet":"# pip install --upgrade google-cloud-texttospeech\nfrom google.cloud import texttospeech\n\nclient = texttospeech.TextToSpeechClient()\nconfig = texttospeech.StreamingSynthesizeConfig(\n    voice=texttospeech.VoiceSelectionParams(\n        name=\"en-US-Chirp3-HD-Charon\", language_code=\"en-US\"))\n\ndef requests():\n    yield texttospeech.StreamingSynthesizeRequest(streaming_config=config)\n    for text in [\"Hello there. \", \"How are you \", \"today?\"]:  # e.g. LLM tokens\n        yield texttospeech.StreamingSynthesizeRequest(\n            input=texttospeech.StreamingSynthesisInput(text=text))\n\nwith open(\"out.pcm\", \"wb\") as f:\n    for resp in client.streaming_synthesize(requests()):\n        f.write(resp.audio_content)  # PCM by default for streaming"},"warnings":[{"severity":"high","title":"Bidi streaming only for Chirp 3 HD and still Preview","detail":"The text-in streaming page says StreamingSynthesize is only compatible with Chirp 3 HD voices and is a Pre-GA feature with limited support. Gemini-TTS streams output but you still send the text up front."},{"severity":"high","title":"Gemini 3.8 TTS is not on the Cloud TTS API","detail":"Gemini 3.8 Flash and Flash-Lite TTS are Preview and only available through the Gemini Enterprise API. Their prices double on Jan 1 2027."},{"severity":"medium","title":"Whitespace and SSML tags are billed","detail":"Google counts every character including spaces, newlines and SSML tags (except <mark>). Strip markdown and extra whitespace from LLM output before sending."},{"severity":"medium","title":"Token-billed Gemini voices scale with audio length","detail":"Gemini-TTS bills 25 audio tokens per second, so slow pacing prompts or long pauses raise cost independent of text length."},{"severity":"medium","title":"gRPC and IAM auth","detail":"Streaming uses gRPC with Google credentials; browsers cannot call it directly. Run a server-side relay."},{"severity":"low","title":"Studio voices are very expensive","detail":"Legacy Studio voices are $160/1M characters; do not pick them for agents."}],"best_for":"GCP shops needing many languages and regional endpoints; Gemini-TTS for prompt-styled or multi-speaker output.","open_source":false,"self_hostable":false,"compliance":"Google Cloud data terms and regional endpoints; certifications not re-verified here.","docs":[{"label":"Bidirectional streaming","url":"https://docs.cloud.google.com/text-to-speech/docs/create-audio-text-streaming"},{"label":"Chirp 3 HD","url":"https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd"},{"label":"Gemini-TTS","url":"https://docs.cloud.google.com/text-to-speech/docs/gemini-tts"}],"sources":["https://cloud.google.com/text-to-speech/pricing","https://docs.cloud.google.com/text-to-speech/docs/create-audio-text-streaming","https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd","https://docs.cloud.google.com/text-to-speech/docs/gemini-tts"],"confidence":"high","unverified":"Latency figures; per-project streaming quotas.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"Google Cloud TTS","facts":{"cat":"tts","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":false,"grpc":true,"self_hostable":false,"open_weights":false,"price_per_1m_chars":30,"voices":30,"voice_cloning":true,"instant_cloning":true,"text_streaming_input":true,"word_timestamps":null,"emotion_control":true,"ssml":true,"telephony_8k":true,"commercial_free_tier":null,"_notes":"Figures are Chirp 3 HD: 30 voices, 53 locales, $30/1M with the first 1M chars/month free. Text-in streaming is Chirp 3 HD only and still Preview. Gemini-TTS is token-billed with prompt-based style control. Instant custom voice $60/1M. SSML on legacy voices only."}},{"id":"azure-speech-tts","name":"Azure AI Speech neural and HD voices","vendor":"Microsoft","category":"tts","summary":"Huge catalog (500+ prebuilt voices) with a WebSocket v2 endpoint that accepts streamed text from an LLM via the Speech SDK. DragonHD voices add emotion-aware expressiveness; MAI-Voice is a newer premium tier.","status":"GA","models":[{"name":"Prebuilt neural voices (e.g. en-US-AvaMultilingualNeural)","status":"GA","notes":"Work with text streaming on the v2 endpoint."},{"name":"DragonHD (e.g. en-US-Ava:DragonHDLatestNeural)","status":"GA / some Preview voices","notes":"30+ fine-tuned HD voices, real-time only, subset of SSML."},{"name":"DragonHDOmni","status":"per docs","notes":"500+ voices with style support."},{"name":"Azure OpenAI voices in Speech","status":"GA","notes":"Not supported by text streaming; >500 ms latency per Microsoft comparison."},{"name":"MAI-Voice","status":"Listed on pricing page","notes":"Price not resolvable from the public page or retail price API at time of research."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text or SSML (text streaming mode does not support SSML)","output":"opus, mp3, pcm, truesilk at 8/16/24/48 kHz; raw PCM formats such as Raw24Khz16BitMonoPcm"},"languages":"Many locales; see language-support page (count not re-verified)","voices":"More than 500 prebuilt; Personal Voice and Professional custom voice (gated access)","latency":"Microsoft comparison table: HD and standard neural voices < 300 ms; Azure OpenAI voices > 500 ms.","features":["input text streaming (Speech SDK, C#/C++/Python)","word boundary events","visemes","custom neural voice","personal voice cloning (gated)","containers / embedded / disconnected deployment for non-HD voices","commitment tiers"],"pricing":{"model":"per-character","items":[{"what":"Neural (real-time and batch)","price":"$15","unit":"per 1M characters","notes":"Azure retail price API, eastus"},{"what":"Neural HD","price":"$22","unit":"per 1M characters","notes":"Azure retail price API"},{"what":"Custom neural (professional) real-time","price":"$24","unit":"per 1M characters","notes":"Plus $4.032/hour endpoint hosting and training at $52/compute hour"},{"what":"Custom neural HD synthesis","price":"$48","unit":"per 1M characters","notes":""},{"what":"Personal Voice synthesis","price":"$24","unit":"per 1M characters","notes":"Gated feature"},{"what":"Commitment tiers (Neural)","price":"$960/80M to $24,000/4,000M per month","unit":"monthly","notes":"Overage $12 down to $6 per 1M"}],"est_per_minute_usd":{"low":0.0135,"high":0.0198,"basis":"900 chars/min; Neural $15/1M vs Neural HD $22/1M at pay-as-you-go"},"free_tier":"F0: 0.5M neural characters per month","source":"https://prices.azure.com/api/retail/prices (filter: Azure Speech, eastus) and https://azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/"},"limits":["Text streaming: SDK only (C#, C++, Python), WebSocket v2 endpoint required","HD voices: real-time only, subset of SSML, cloud only","Concurrency per resource defaults; see quotas page"],"regions":"Dozens of Azure regions for standard neural voices; HD voices in a subset","setup":{"steps":["Create a Speech resource in the Azure portal; copy key and region.","pip install azure-cognitiveservices-speech.","Point SpeechConfig at wss://{region}.tts.speech.microsoft.com/cognitiveservices/websocket/v2.","Create a SpeechSynthesisRequest with input_type TextStream and write LLM chunks into input_stream, then close it."],"endpoint":"wss://{region}.tts.speech.microsoft.com/cognitiveservices/websocket/v2","auth":"Speech resource key (subscription) or Entra ID token","snippet_lang":"python","snippet":"# pip install azure-cognitiveservices-speech\nimport os\nimport azure.cognitiveservices.speech as speechsdk\n\nendpoint = f\"wss://{os.environ['AZURE_TTS_REGION']}.tts.speech.microsoft.com/cognitiveservices/websocket/v2\"\ncfg = speechsdk.SpeechConfig(endpoint=endpoint, subscription=os.environ[\"AZURE_TTS_API_KEY\"])\ncfg.speech_synthesis_voice_name = \"en-US-AvaMultilingualNeural\"\n\nsynth = speechsdk.SpeechSynthesizer(speech_config=cfg)  # default speaker output\nreq = speechsdk.SpeechSynthesisRequest(\n    input_type=speechsdk.SpeechSynthesisRequestInputType.TextStream)\ntask = synth.speak_async(req)\n\nfor chunk in [\"Hello there. \", \"This text arrives \", \"from an LLM stream.\"]:\n    req.input_stream.write(chunk)\nreq.input_stream.close()\n\nresult = task.get()\nprint(result.reason)"},"warnings":[{"severity":"high","title":"Text streaming needs the SDK and the v2 endpoint","detail":"Text-in streaming only works through the Speech SDK (C#, C++, Python) against /cognitiveservices/websocket/v2. Plain REST or the v1 socket will not accept partial text. No JavaScript support listed."},{"severity":"medium","title":"No SSML in text streaming","detail":"When streaming text you set voice and format as global properties; SSML (prosody, breaks, styles) is not supported in that mode."},{"severity":"medium","title":"Azure OpenAI voices excluded","detail":"OpenAI voices inside Azure Speech are not supported by text streaming and Microsoft lists them at >500 ms latency."},{"severity":"medium","title":"Custom voice hosting is billed hourly","detail":"Professional custom voices cost $4.032 per model per hour to host on top of per-character synthesis; an idle endpoint still costs about $2,900/month."},{"severity":"low","title":"Pricing page renders without numbers","detail":"The public pricing page loads prices dynamically; use the Azure retail price API or calculator to confirm your region."}],"best_for":"Enterprises on Azure needing many languages/voices, private networking, containers or disconnected deployment.","open_source":false,"self_hostable":true,"compliance":"Azure compliance programs; containers and disconnected options for non-HD voices. Specific certifications not re-verified here.","docs":[{"label":"Lower synthesis latency / text streaming","url":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/how-to-lower-speech-synthesis-latency"},{"label":"HD voices","url":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/high-definition-voices"}],"sources":["https://prices.azure.com/api/retail/prices","https://azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/","https://learn.microsoft.com/en-us/azure/ai-services/speech-service/how-to-lower-speech-synthesis-latency","https://learn.microsoft.com/en-us/azure/ai-services/speech-service/high-definition-voices"],"confidence":"high","unverified":"MAI-Voice and 'Neural HD Flash' prices (shown as categories on the pricing page but not resolvable); exact locale count.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"Azure Speech TTS","facts":{"cat":"tts","latency_ms":300,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_1m_chars":15,"voices":500,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":true,"emotion_control":true,"ssml":true,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Microsoft lists neural and HD voices under 300 ms. $15/1M Neural, $22/1M Neural HD (eastus). Voices: 500+ (DragonHDOmni). Text streaming only via the Speech SDK (C#, C++, Python) on the v2 endpoint, without SSML. Free tier 0.5M chars/month. Cloning via custom neural and gated personal voice. Compliance per Azure scope, not re-verified."}},{"id":"amazon-polly","name":"Amazon Polly (generative and bidirectional streaming)","vendor":"Amazon Web Services","category":"tts","summary":"Polly now has StartSpeechSynthesisStream, an HTTP/2 bidirectional API that takes text incrementally and returns audio as it is produced, but only for the generative engine. Standard and neural engines remain request/response with streamed output.","status":"GA","models":[{"name":"generative engine","status":"GA","notes":"Only engine accepted by StartSpeechSynthesisStream."},{"name":"long-form engine","status":"GA","notes":"$100/1M chars; not for real-time."},{"name":"neural engine","status":"GA","notes":"SynthesizeSpeech with streamed output."},{"name":"standard engine","status":"GA","notes":"Cheapest, oldest."}],"transports":["HTTP chunked","HTTP/2 bidirectional event stream"],"audio":{"input":"Plain text or SSML","output":"mp3, ogg_opus, ogg_vorbis, pcm (stream API; JSON speech marks not supported on the stream API)"},"languages":"Stream API LanguageCode list covers ~40 locales (only needed for bilingual voices)","voices":"Generative voices subset of Polly catalog; no voice cloning","latency":"No numeric claim found.","features":["bidirectional text-in streaming (generative only)","flush via FlushStreamConfiguration","lexicons (up to 5)","speech marks (non-stream API)"],"pricing":{"model":"per-character","items":[{"what":"Generative voices","price":"$30","unit":"per 1M characters","notes":""},{"what":"Neural voices","price":"$16","unit":"per 1M characters","notes":""},{"what":"Standard voices","price":"$4","unit":"per 1M characters","notes":""},{"what":"Long-form voices","price":"$100","unit":"per 1M characters","notes":""}],"est_per_minute_usd":{"low":0.0144,"high":0.027,"basis":"900 chars/min. Bidi streaming requires generative ($30/1M = $0.027); neural $16/1M = $0.0144 without input streaming."},"free_tier":"12 months: 5M standard, 1M neural, 500K long-form, 100K generative characters per month; new accounts also get up to $200 AWS credits","source":"https://aws.amazon.com/polly/pricing/"},"limits":["Bidi stream: generative engine only","Up to 5 lexicons","GovCloud prices differ ($4.80 standard, $19.20 neural)"],"regions":"Bidi streaming region list not confirmed officially; a third-party package lists us-east-1, us-west-2, eu-central-1, eu-west-2, ap-southeast-1, ca-central-1 (2026-05)","setup":{"steps":["Create IAM credentials with polly:SynthesizeSpeech (and the stream action).","For true text-in streaming use an SDK that exposes StartSpeechSynthesisStream (AWS Java SDK per third-party notes; boto3 support unconfirmed).","Otherwise call SynthesizeSpeech per sentence with Engine=generative and read the AudioStream."],"endpoint":"POST /v1/synthesisStream (StartSpeechSynthesisStream); SynthesizeSpeech for request/response","auth":"AWS SigV4 (IAM credentials)","snippet_lang":"python","snippet":"# pip install boto3  - per-sentence request with streamed output\nimport boto3\n\npolly = boto3.client(\"polly\", region_name=\"us-east-1\")\nresp = polly.synthesize_speech(\n    Text=\"Hello from Amazon Polly's generative engine.\",\n    VoiceId=\"Ruth\",            # pick a voice that supports the generative engine\n    Engine=\"generative\",\n    OutputFormat=\"pcm\",\n    SampleRate=\"16000\",\n)\nwith open(\"out_16k_s16le.pcm\", \"wb\") as f:\n    for chunk in resp[\"AudioStream\"].iter_chunks():\n        f.write(chunk)\n# For text-in streaming use StartSpeechSynthesisStream (generative only)\n# from an SDK that supports HTTP/2 event streams."},"warnings":[{"severity":"high","title":"Bidi streaming is generative-only","detail":"StartSpeechSynthesisStream accepts only the generative engine even though the parameter lists others. Generative costs $30/1M chars, nearly 2x neural."},{"severity":"medium","title":"SDK support is uneven","detail":"A third-party package says boto3 does not expose StartSpeechSynthesisStream and only the Java SDK does. Check your SDK version before designing around it."},{"severity":"medium","title":"Generative free tier is tiny","detail":"Only 100K generative characters/month for 12 months, about 2 hours of audio."},{"severity":"low","title":"No voice cloning","detail":"Polly has no self-serve cloning; brand voices are an enterprise engagement."},{"severity":"low","title":"Speech marks not on the stream API","detail":"JSON speech marks (word timings) are not supported by the bidi stream; use separate SynthesizeSpeech calls if you need them."}],"best_for":"AWS-native stacks (Connect, Lex) wanting IAM auth and predictable per-character pricing.","open_source":false,"self_hostable":false,"compliance":"AWS compliance programs (GovCloud availability for standard/neural).","docs":[{"label":"StartSpeechSynthesisStream","url":"https://docs.aws.amazon.com/polly/latest/dg/API_StartSpeechSynthesisStream.html"},{"label":"Pricing","url":"https://aws.amazon.com/polly/pricing/"}],"sources":["https://aws.amazon.com/polly/pricing/","https://docs.aws.amazon.com/polly/latest/dg/API_StartSpeechSynthesisStream.html","https://amazon-polly-streaming.readthedocs.io/en/stable/overview.html"],"confidence":"medium","unverified":"Which SDKs support the bidi stream; region list; whether 'Ruth' is the right generative voice for your locale.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"Amazon Polly","facts":{"cat":"tts","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":200,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":30,"voices":null,"voice_cloning":false,"instant_cloning":false,"text_streaming_input":true,"word_timestamps":true,"emotion_control":null,"ssml":true,"telephony_8k":false,"commercial_free_tier":null,"_notes":"$30/1M is the generative engine, the only one that accepts bidirectional text streaming (neural $16/1M). Transport is an HTTP/2 event stream. Stream API covers about 40 locales. Speech marks (word timings) only on the non-stream API. Free tier is 100K generative chars/month for 12 months; up to $200 AWS credits for new accounts. EU region per third-party list."}},{"id":"rime","name":"Rime TTS (Coda, Mist v3)","vendor":"Rime Labs","category":"tts","summary":"Conversational TTS aimed at contact-centre agents, with a JSON WebSocket (/ws3), word timestamps, native mu-law and an on-prem option. Arcana was retired from the cloud on 2026-08-15 and now routes to Coda.","status":"GA","models":[{"name":"coda","status":"GA (launched 2026-05-19)","notes":"Successor to Arcana; English, French, German, Japanese, Portuguese, Spanish, plus Arabic and Hindi (2026-08-04)."},{"name":"mistv3","status":"GA (2026-04)","notes":"Lowest cloud latency; English, French, German, Spanish."},{"name":"mistv2 / mistv1","status":"Legacy","notes":"Still served; /ws2 limited to these."},{"name":"arcana / arcanav2 / arcanav3","status":"Retired from cloud 2026-08-15","notes":"Requests now route to Coda."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text (custom pronunciation via spell() and dictionaries)","output":"pcm, mp3, mulaw (audioFormat); sample rates per model reference"},"languages":"Coda: 8; Mist v3: 4","voices":"Catalog voices; custom voice clones on Enterprise","latency":"Vendor: sub-200 ms end-to-end via cloud API; Mist v3 typical TTFB well below 100 ms; Coda sub-100 ms model latency on GPU engine (self-hosted) plus 25-50 ms network from most of the US.","features":["input streaming (/ws3)","word timestamps","contextId echo","flush / clear / eos operations","segment modes (never, bySentence, immediate)","on-prem Docker/Kubernetes","MCP server"],"pricing":{"model":"per-character","items":[{"what":"Mist v3","price":"$0.03","unit":"per 1K characters","notes":"Starter; vendor says ~$0.03/minute"},{"what":"Coda","price":"$0.05","unit":"per 1K characters","notes":"Starter; ~$0.05/minute"},{"what":"Enterprise","price":"Custom","unit":"","notes":"Unlimited concurrency, BAA, SOC 2 reports, on-prem/VPC"}],"est_per_minute_usd":{"low":0.027,"high":0.045,"basis":"900 chars/min; Mist v3 vs Coda at Starter rates"},"free_tier":"Page conflicts: '~800 minutes free (about 800k characters)' vs FAQ '3,000 free minutes'","source":"https://www.rime.ai/pricing"},"limits":["Starter: 20 concurrent TTS generations","Rime keeps only one active contextId at a time per socket"],"regions":"users-ws.rime.ai plus users-east-ws.rime.ai (us-east-1) and users-west regional hosts","setup":{"steps":["Create an API key at rime.ai.","Connect to wss://users-ws.rime.ai/ws3 with speaker, modelId and audioFormat query params and an Authorization: Bearer header.","Send {text} messages, then {operation:'flush'} or {operation:'eos'}.","Decode base64 'chunk' events; use 'timestamps' for barge-in."],"endpoint":"wss://users-ws.rime.ai/ws3","auth":"Authorization: Bearer <RIME_API_KEY> header (browsers cannot set it; proxy via backend)","snippet_lang":"javascript","snippet":"// npm i ws\nimport WebSocket from \"ws\";\nimport fs from \"fs\";\n\nconst url = \"wss://users-ws.rime.ai/ws3?speaker=YOUR_SPEAKER&modelId=mistv3&audioFormat=pcm\";\nconst ws = new WebSocket(url, { headers: { Authorization: `Bearer ${process.env.RIME_API_KEY}` } });\nconst out = fs.createWriteStream(\"out.pcm\");\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({ text: \"Hello there. \", contextId: \"turn-1\" }));\n  ws.send(JSON.stringify({ text: \"Thanks for calling today.\" }));\n  ws.send(JSON.stringify({ operation: \"eos\" }));   // speak the rest, then close\n});\nws.on(\"message\", (raw) => {\n  const m = JSON.parse(raw.toString());\n  if (m.type === \"chunk\") out.write(Buffer.from(m.data, \"base64\"));\n  if (m.type === \"error\") console.error(m);\n});\nws.on(\"close\", () => out.end());"},"warnings":[{"severity":"high","title":"Arcana is gone from the cloud","detail":"Since 2026-08-15 arcana, arcanav2 and arcanav3 requests are silently routed to Coda. Voices and pricing may differ from what you tested; re-evaluate."},{"severity":"medium","title":"Endpoint/model matrix is inconsistent in docs","detail":"One reference lists mistv2 as not supported on /ws3 while the overview lists it. Test your exact modelId on /ws3 before shipping."},{"severity":"medium","title":"Free allowance is unclear","detail":"The pricing page says ~800 minutes free while the FAQ says 3,000 minutes. Do not plan around either number."},{"severity":"low","title":"Single context per socket","detail":"Rime does not track multiple simultaneous contextIds; open one socket per concurrent speaker."},{"severity":"low","title":"Per-minute price assumes ~1K chars/min","detail":"Rime equates $0.03/1K chars with ~$0.03/min; dense text costs more."}],"best_for":"US contact-centre and telephony agents that need natural conversational English, mu-law output and an on-prem path.","open_source":false,"self_hostable":true,"compliance":"Enterprise: BAA and SOC 2 reports; optional request-data retention controls (disabled by default).","docs":[{"label":"WebSockets overview","url":"https://docs.rime.ai/docs/websockets"},{"label":"Changelog","url":"https://docs.rime.ai/docs/changelog"},{"label":"JSON WebSocket reference","url":"https://docs.rime.ai/api-reference/websockets-json"}],"sources":["https://www.rime.ai/pricing","https://docs.rime.ai/docs/websockets","https://docs.rime.ai/docs/changelog","https://docs.rime.ai/docs/latency"],"confidence":"high","unverified":"Per-model sample rates; speaker names for mistv3 (placeholder in snippet).","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":200,"languages":8,"max_session_min":null,"concurrency":20,"free_tier":true,"free_credit_usd":null,"hipaa":true,"soc2":true,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":false,"price_per_1m_chars":50,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":true,"emotion_control":null,"ssml":null,"telephony_8k":true,"commercial_free_tier":null,"_notes":"Figures are Coda ($50/1M, 8 languages); Mist v3 is $30/1M with 4 languages and TTFB well below 100 ms. Cloud claim sub-200 ms end to end. Concurrency 20 on Starter. BAA and SOC 2 reports on Enterprise. Free allowance unclear (800 vs 3,000 minutes)."}},{"id":"hume-octave","name":"Hume Octave TTS","vendor":"Hume AI","category":"tts","summary":"LLM-based expressive TTS with voice design from text prompts. Hume is sunsetting its TTS and EVI APIs: access ends 2026-11-13 and account data is deleted afterwards.","status":"Deprecated","models":[{"name":"Octave 1","status":"Sunsetting 2026-11-13","notes":"English and Spanish, ~200 ms (vendor)."},{"name":"Octave 2 (preview)","status":"Sunsetting 2026-11-13","notes":"11 languages, ~100 ms model latency excluding network (vendor)."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text plus optional acting description","output":"MP3, WAV, PCM"},"languages":"Octave 2: English, Japanese, Korean, Spanish, French, Portuguese, Italian, German, Russian, Hindi, Arabic","voices":"Voice library, voice design, cloning from ~15 s audio","latency":"Vendor: Octave 2 ~100 ms model latency; instant mode first audio ~200 ms.","features":["input streaming (/v0/tts/stream/input)","voice design from text prompts","voice cloning","instant mode"],"pricing":{"model":"subscription","items":[{"what":"Pricing","price":"Not published on hume.ai/pricing at time of research","unit":"","notes":"Product is being shut down"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Not verifiable; API ends 2026-11-13"},"free_tier":"Not verified","source":"https://www.hume.ai/pricing"},"limits":["5,000 characters per utterance","1,000-character descriptions","Up to 5 generations per request"],"regions":"Not specified","setup":{"steps":["Do not start new integrations: access ends 2026-11-13.","Existing users: export anything you need and migrate before that date."],"endpoint":"wss://api.hume.ai/v0/tts/stream/input (path per docs; host not re-verified)","auth":"X-Hume-Api-Key header (not re-verified)","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"API shutdown on 2026-11-13","detail":"Hume's docs state TTS and EVI access ends November 13, 2026 at 12:01 a.m. EST and account data is permanently deleted after that date. Migrate now."},{"severity":"medium","title":"Data deletion","detail":"Cloned and designed voices stored in your Hume account will be deleted; there is no indication they can be exported to another vendor."},{"severity":"low","title":"Instant mode restrictions","detail":"Instant mode needs a predefined voice and num_generations of 1; voice design requests are slower."}],"best_for":"Nothing new; existing users should migrate (e.g. to Inworld TTS-2 or ElevenLabs v4 for prompt-steered expressiveness).","open_source":false,"self_hostable":false,"compliance":"N/A (sunsetting)","docs":[{"label":"TTS overview (sunset notice)","url":"https://dev.hume.ai/docs/text-to-speech-tts/overview"}],"sources":["https://dev.hume.ai/docs/text-to-speech-tts/overview","https://www.hume.ai/pricing"],"confidence":"high","unverified":"Pricing; exact WebSocket host/auth header.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":100,"languages":11,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":null,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":null,"emotion_control":true,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"API shuts down 2026-11-13 and account data is deleted. Octave 2 (preview): ~100 ms model latency, 11 languages. Pricing not published."}},{"id":"lmnt","name":"LMNT","vendor":"LMNT","category":"tts","summary":"Former low-latency TTS and voice-cloning API. The lmnt.com site now states 'LMNT has shut down'.","status":"Shut down","models":[{"name":"aurora / blizzard","status":"Shut down","notes":"Historical model names; service discontinued."}],"transports":[],"audio":{"input":"N/A","output":"N/A"},"languages":"N/A","voices":"N/A","latency":"N/A","features":[],"pricing":{"model":"subscription","items":[],"est_per_minute_usd":{"low":null,"high":null,"basis":"Service shut down"},"free_tier":"N/A","source":"https://lmnt.com/"},"limits":[],"regions":"N/A","setup":{"steps":["Remove LMNT from your stack; choose another provider."],"endpoint":"N/A","auth":"N/A","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"Service has shut down","detail":"lmnt.com and lmnt.com/pricing both say 'LMNT has shut down. Our speech generation journey has come to an end.' Third-party directories and SDK packages still list it as active; ignore them."},{"severity":"medium","title":"Stale SDKs and integrations","detail":"Packages and framework plugins for LMNT may still install without error but will not work. Remove them to avoid runtime failures."},{"severity":"low","title":"Pick a replacement by the feature you used","detail":"LMNT was mainly used for fast streaming and quick cloning; Cartesia, Inworld (TTS-2 Flash) or self-hosted Chatterbox cover the same ground."}],"best_for":"N/A","open_source":false,"self_hostable":false,"compliance":"N/A","docs":[{"label":"Shutdown notice","url":"https://lmnt.com/"}],"sources":["https://lmnt.com/","https://lmnt.com/pricing"],"confidence":"high","unverified":"Exact shutdown date (not stated on the site).","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"grpc":null,"self_hostable":false,"open_weights":false,"price_per_1m_chars":null,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":null,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Service has shut down."}},{"id":"inworld-tts","name":"Inworld TTS (Realtime TTS-2, TTS-2 Flash)","vendor":"Inworld AI","category":"tts","summary":"Low-cost, fast TTS with natural-language steering on TTS-2 and a very fast Flash variant. Bidirectional WebSocket with multiple contexts, timestamps, cloning and an OpenAI-compatible endpoint.","status":"GA","models":[{"name":"inworld-tts-2","status":"GA (2026-05-05)","notes":"Steering via natural language and inline tags; P90 server-side TTFB 100 ms (vendor)."},{"name":"inworld-tts-2-flash","status":"GA (2026-08-09)","notes":"P90 server-side TTFB 20 ms (vendor); steering ignored; cheapest."},{"name":"inworld-tts-1.5-max / inworld-tts-1.5-mini","status":"Legacy (Jan 2026)","notes":"Still usable per integration docs."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text with inline tags (pauses, pronunciation, steering on TTS-2)","output":"LINEAR16 (WAV header per chunk), PCM, MP3 (default), OGG_OPUS, ALAW, MULAW, WAV; 8-48 kHz (default 48 kHz)"},"languages":"Docs say 200+ languages and locales; release notes describe 15 production-quality plus 90+ experimental for TTS-2","voices":"Catalog voices; instant cloning on all plans (100 custom voices on On-Demand); professional cloning (beta) on TTS-2","latency":"Vendor: P90 server-side TTFB 100 ms (TTS-2), 20 ms (TTS-2 Flash).","features":["input streaming (bidirectional WebSocket)","multiple contexts per connection","autoMode sentence buffering","word timestamps, phonemes and visemes (TTS-2)","voice cloning","voice design","zero data retention","OpenAI-compatible POST /v1/audio/speech (2026-09-11)"],"pricing":{"model":"per-character","items":[{"what":"TTS-2 Flash, On-Demand","price":"$15","unit":"per 1M characters","notes":"Free to start, up to 70 min TTS included"},{"what":"TTS-2, On-Demand","price":"$25","unit":"per 1M characters","notes":""},{"what":"Creator $25/mo","price":"$20 / $10","unit":"per 1M characters (TTS-2 / Flash)","notes":"Concurrency 10"},{"what":"Builder $100/mo","price":"$17.50 / $9","unit":"per 1M characters","notes":"Concurrency 50"},{"what":"Developer $300/mo","price":"$15 / $8","unit":"per 1M characters","notes":"Concurrency 150"},{"what":"Growth $1,500/mo","price":"$12.50 / $7","unit":"per 1M characters","notes":"Concurrency 500"},{"what":"Enterprise","price":"As low as $5 (TTS-2), sub-$5 (Flash)","unit":"per 1M characters","notes":"Custom"}],"est_per_minute_usd":{"low":0.0063,"high":0.0225,"basis":"900 chars/min. Low = TTS-2 Flash on Growth ($7/1M); high = TTS-2 on On-Demand ($25/1M)."},"free_tier":"On-Demand: up to 70 minutes of TTS","source":"https://inworld.ai/pricing"},"limits":["Concurrent requests: On-Demand 5, Creator 10, Builder 50, Developer 150, Growth 500","2,000 UTF-16 code units per send_text message","Socket closes after 10 min of inactivity across all contexts","Sync 2,000 chars, HTTP streaming 4,000 chars, async 100,000 chars per request"],"regions":"Not specified","setup":{"steps":["Get an API key (Basic credential) from the Inworld portal.","Connect to wss://api.inworld.ai/tts/v1/voice:streamBidirectional with Authorization: Basic <key>.","Send create (voice_id, model_id, audio_config), then send_text messages, then flush_context / close_context.","Decode result.audioChunk.audioContent (base64)."],"endpoint":"wss://api.inworld.ai/tts/v1/voice:streamBidirectional","auth":"Authorization: Basic <INWORLD_API_KEY> (the key from the portal is already Base64)","snippet_lang":"javascript","snippet":"// npm i ws  - message shapes follow Inworld's official example_websocket.js\nimport WebSocket from \"ws\";\nimport fs from \"fs\";\n\nconst ws = new WebSocket(\"wss://api.inworld.ai/tts/v1/voice:streamBidirectional\", {\n  headers: { Authorization: `Basic ${process.env.INWORLD_API_KEY}` },\n});\nconst out = fs.createWriteStream(\"out_24k_s16le.pcm\");\nconst ctx = \"turn-1\";\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({ context_id: ctx, create: {\n    voice_id: \"Ashley\", model_id: \"inworld-tts-2-flash\",\n    audio_config: { audio_encoding: \"PCM\", sample_rate_hertz: 24000 } } }));\n  for (const t of [\"Hello there. \", \"Streaming text \", \"from an LLM.\"])\n    ws.send(JSON.stringify({ context_id: ctx, send_text: { text: t } }));\n  ws.send(JSON.stringify({ context_id: ctx, close_context: {} }));\n});\nws.on(\"message\", (raw) => {\n  const r = JSON.parse(raw.toString()).result;\n  const b64 = r?.audioChunk?.audioContent || r?.audioContent;\n  if (b64) out.write(Buffer.from(b64, \"base64\"));\n  if (r?.contextClosed) ws.close();\n});\nws.on(\"close\", () => out.end());"},"warnings":[{"severity":"medium","title":"Steering is ignored on Flash","detail":"inworld-tts-2-flash ignores the instruction field and [shouting]-style tags. Use inworld-tts-2 if you rely on prompt-based delivery."},{"severity":"medium","title":"LINEAR16 puts a WAV header in every chunk","detail":"With audio_encoding LINEAR16 each chunk carries a WAV header, which clicks if concatenated. Use PCM (headerless) for streaming playback."},{"severity":"medium","title":"Characters counted in UTF-16 code units","detail":"Limits and billing count UTF-16 code units, so many emoji and some scripts count as two. Strip emoji from LLM output."},{"severity":"medium","title":"Language claims vary","detail":"Docs say 200+ languages while release notes say 15 production-quality plus 90+ experimental. Test non-English quality before committing."},{"severity":"low","title":"Steering persistence changed","detail":"Since 2026-08-06 steering instructions persist until changed or [reset]; older code that assumed per-request tags may behave differently."}],"best_for":"Cost-sensitive, high-volume voice agents and games that still want cloning and timestamps.","open_source":false,"self_hostable":false,"compliance":"Zero data retention supported (models page). Certifications not re-verified.","docs":[{"label":"TTS models","url":"https://docs.inworld.ai/tts/tts-models.md"},{"label":"WebSocket API reference","url":"https://docs.inworld.ai/api-reference/ttsAPI/texttospeech/synthesize-speech-websocket.md"},{"label":"Release notes","url":"https://docs.inworld.ai/release-notes/tts"},{"label":"Official WebSocket example","url":"https://github.com/inworld-ai/inworld-api-examples/blob/main/tts/js/example_websocket.js"}],"sources":["https://inworld.ai/pricing","https://docs.inworld.ai/tts/tts-models.md","https://docs.inworld.ai/release-notes/tts","https://docs.inworld.ai/api-reference/ttsAPI/texttospeech/synthesize-speech-websocket.md","https://docs.inworld.ai/tts/synthesize-speech.md","https://github.com/inworld-ai/inworld-api-examples/blob/main/tts/js/example_websocket.js"],"confidence":"high","unverified":"Voice name 'Ashley' availability on TTS-2 Flash; per-plan WebSocket connection caps.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"Inworld TTS","facts":{"cat":"tts","latency_ms":100,"languages":15,"max_session_min":null,"concurrency":5,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":25,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":true,"emotion_control":true,"ssml":null,"telephony_8k":true,"commercial_free_tier":null,"_notes":"Figures are TTS-2 (P90 server TTFB 100 ms, $25/1M On-Demand). TTS-2 Flash: 20 ms P90, $15/1M, ignores steering. Languages: 15 production quality per release notes; docs claim 200+. Concurrency 5 On-Demand (Creator 10, Builder 50). Free start includes up to 70 minutes. Zero data retention supported."}},{"id":"fish-audio","name":"Fish Audio API (S2.1 Pro)","vendor":"Fish Audio","category":"tts","summary":"Multilingual (83 languages per third parties) cloning-first TTS with a MessagePack WebSocket for streamed text. Billed per UTF-8 byte, with a free s2.1-pro-free model under fair use.","status":"GA","models":[{"name":"s2.1-pro","status":"GA (default if model header omitted)","notes":"$15/M UTF-8 bytes."},{"name":"s2.1-pro-free","status":"Free, fair use, no SLA","notes":"Third-party sources say available through 2026-11-30; requests may be retained for training."},{"name":"s2-pro","status":"GA","notes":"Multi-speaker dialogue supported."},{"name":"s1","status":"Deprecated, retiring 2026-12-31","notes":"After that, s1 requests are served and billed as s2.1-pro."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text (MessagePack frames on WebSocket)","output":"mp3 (default, 64/128/192 kbps), wav, pcm, opus; 44.1 kHz for most formats, 48 kHz for opus"},"languages":"83 per third-party coverage of S2.1 Pro (not on the docs pages checked)","voices":"Very large community voice library; instant cloning via references","latency":"Vendor claim (third-party reported): ~90 ms to first audio for S2.1 Pro; Vapi measured 141 ms median including network.","features":["input streaming (MessagePack WebSocket)","latency modes low/balanced/normal","voice cloning from reference audio","multi-speaker dialogue","prosody speed/volume"],"pricing":{"model":"per-character","items":[{"what":"s2.1-pro / s2-pro / s1","price":"$15.00","unit":"per 1M UTF-8 bytes","notes":"No subscription or minimum for API"},{"what":"s2.1-pro-free","price":"$0.00","unit":"per 1M UTF-8 bytes","notes":"Fair use, no SLA"},{"what":"voice-design-1","price":"$0.01","unit":"per successful request","notes":""}],"est_per_minute_usd":{"low":0.0135,"high":0.0405,"basis":"900 chars/min. English ASCII = 1 byte/char ($0.0135/min); CJK/Arabic/Hindi are ~3 bytes/char (~$0.04/min). Free model excluded."},"free_tier":"s2.1-pro-free model at $0 under fair use (time-limited per third-party sources)","source":"https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits.md"},"limits":["Concurrent requests: Starter (<$100 paid) 5, Elevated ($100+) 15, High Volume ($1,000+) 50","429 without Retry-After; use exponential backoff","chunk_length 100-300"],"regions":"Not specified","setup":{"steps":["Create an API key and prepay credit at fish.audio.","Choose a voice reference_id (or upload references).","Open wss://api.fish.audio/v1/tts/live with Authorization: Bearer and a model header.","Send start, text..., flush, stop as MessagePack; collect 'audio' events until 'finish'.","Or use the Python SDK's stream_websocket with a text generator."],"endpoint":"wss://api.fish.audio/v1/tts/live","auth":"Authorization: Bearer <FISH_API_KEY>; header model: s2.1-pro","snippet_lang":"python","snippet":"# Fish Audio Python SDK (import name 'fishaudio'; check PyPI for the package name)\nfrom fishaudio import FishAudio\n\nclient = FishAudio()  # reads FISH_API_KEY\n\ndef llm_tokens():\n    for token in [\"The \", \"first \", \"move \", \"sets \", \"everything \", \"in \", \"motion.\"]:\n        yield token\n\nwith open(\"out.mp3\", \"wb\") as f:\n    for chunk in client.tts.stream_websocket(llm_tokens(), reference_id=\"YOUR_VOICE_ID\"):\n        f.write(chunk)  # play or forward as it arrives"},"warnings":[{"severity":"high","title":"Billed per UTF-8 byte, not character","detail":"Non-Latin scripts cost ~3x per character (Chinese, Japanese, Korean, Arabic, Hindi are 3-4 bytes each). Compare vendors on bytes for non-English workloads."},{"severity":"medium","title":"Free model is temporary and may train on your data","detail":"s2.1-pro-free has no SLA, is reported to run only through 2026-11-30, and Fish's blog says requests may be retained for model improvement. Do not send sensitive text."},{"severity":"medium","title":"s1 retires 2026-12-31","detail":"After that date s1 requests are served by s2.1-pro; voices may sound different."},{"severity":"medium","title":"Low default concurrency","detail":"Only 5 concurrent requests until you have paid $100; 429s carry no Retry-After header."},{"severity":"medium","title":"Cloning community voices","detail":"The public voice library contains user-uploaded clones; confirm you have rights to any voice you use commercially."},{"severity":"low","title":"Open weights are non-commercial","detail":"The S2 Pro weights on Hugging Face use the Fish Audio Research License, not a commercial licence; self-hosting for a product needs a separate agreement."}],"best_for":"Multilingual cloning and character voices at a low flat price.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Live TTS WebSocket","url":"https://docs.fish.audio/api-reference/endpoint/websocket/tts-live.md"},{"label":"Pricing and rate limits","url":"https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits.md"},{"label":"Python WebSocket guide","url":"https://docs.fish.audio/developer-guide/sdk-guide/python/websocket.md"}],"sources":["https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits.md","https://docs.fish.audio/api-reference/endpoint/websocket/tts-live.md","https://fish.audio/ko/blog/s2-1-pro-free-api/","https://humannessindex.vapi.ai/models/fish-s2-1-pro","https://huggingface.co/fishaudio/s2-pro"],"confidence":"medium","unverified":"Language count (83) and ~90 ms latency come from third-party coverage; end date of the free model; Python package name.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":90,"languages":83,"max_session_min":null,"concurrency":5,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":true,"price_per_1m_chars":15,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":false,"commercial_free_tier":null,"_notes":"$15 per 1M UTF-8 bytes, so non-Latin scripts cost about 3x per character. Latency (~90 ms) and 83 languages come from third-party coverage. Free model s2.1-pro-free has no SLA, is time-limited and may retain data. Open weights (S2 Pro) are non-commercial. Concurrency 5 until $100 paid."}},{"id":"resemble-ai","name":"Resemble AI TTS (Chatterbox models)","vendor":"Resemble AI","category":"tts","summary":"Voice cloning company whose hosted API serves its Chatterbox models (e.g. chatterbox-turbo). Its WebSocket streams audio out but takes a whole text/SSML payload per request. The public pricing page now covers deepfake detection only.","status":"GA","models":[{"name":"chatterbox-turbo","status":"GA","notes":"Used in the WebSocket docs example; English; vendor says sub-200 ms."},{"name":"chatterbox (multilingual)","status":"GA","notes":"Open-weight family also available (MIT)."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text or SSML, max 3,000 characters excluding tags","output":"wav/pcm etc.; example uses wav 32 kHz PCM_32; JSON (base64) or binary frames"},"languages":"English for Turbo; Chatterbox Multilingual covers 23+ languages","voices":"Voice cloning (rapid and professional)","latency":"Vendor (Chatterbox README): hosted service sub-200 ms.","features":["audio-out streaming over WebSocket","character and phoneme timestamps","voice cloning","PerTh watermarking","deepfake detection products"],"pricing":{"model":"per-minute","items":[{"what":"TTS","price":"Not published on resemble.ai/pricing (detection-only page)","unit":"","notes":"Third-party figures conflict: $0.0005/s (Flex) vs $0.006/s (older PAYG)"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Not verifiable from official pages; third-party sources range from ~$0.03 to ~$0.36 per minute"},"free_tier":"Flex pay-as-you-go credits never expire (detection page); TTS free tier not verified","source":"https://www.resemble.ai/pricing/"},"limits":["WebSocket available on Business plans and above","Default 20 simultaneous sessions cluster-wide and 20 parallel connections per key","3,000 characters per request"],"regions":"Not specified","setup":{"steps":["Get a Business-tier (or higher) API key.","Create or pick a voice and project.","Open wss://websocket.cluster.resemble.ai/stream and send one JSON job per utterance.","Read audio frames until audio_end."],"endpoint":"wss://websocket.cluster.resemble.ai/stream","auth":"API key (Bearer) per docs","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"No incremental text input","detail":"The WebSocket takes a complete text/SSML payload per request and streams audio back. For LLM output you must chunk sentences yourself."},{"severity":"high","title":"WebSocket requires Business plan","detail":"Lower plans get Unauthorized on the streaming socket."},{"severity":"medium","title":"TTS pricing is not public","detail":"resemble.ai/pricing lists only detection products. Get a written quote before committing."},{"severity":"low","title":"Watermarked output","detail":"Chatterbox output carries Resemble's imperceptible PerTh watermark; fine for most uses but note it if you post-process audio."}],"best_for":"Teams that want hosted Chatterbox with cloning plus deepfake detection/watermarking from one vendor.","open_source":false,"self_hostable":true,"compliance":"Enterprise on-prem available; SOC 2 documentation mentioned for enterprise on the pricing page.","docs":[{"label":"WebSocket streaming","url":"https://docs.resemble.ai/voice-generation/text-to-speech/streaming-websocket"}],"sources":["https://www.resemble.ai/pricing/","https://docs.resemble.ai/voice-generation/text-to-speech/streaming-websocket","https://www.voiceflow.com/blog/resemble-ai.md","https://github.com/resemble-ai/chatterbox"],"confidence":"low","unverified":"TTS prices; auth header specifics; full model list.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":200,"languages":1,"max_session_min":null,"concurrency":20,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":true,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":true,"price_per_1m_chars":null,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":false,"word_timestamps":true,"emotion_control":null,"ssml":true,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Chatterbox Turbo (English) sub-200 ms vendor claim; Chatterbox Multilingual open weights cover 23+ languages. TTS pricing not published. WebSocket needs the Business plan; 20 sessions is the default there. Timestamps are character and phoneme level. SOC 2 documentation mentioned for enterprise."}},{"id":"murf-falcon","name":"Murf Falcon / Falcon 2","vendor":"Murf AI","category":"tts","summary":"Budget real-time TTS marketed at 1 cent per minute with sub-100 ms time-to-first-audio for Falcon 2, geo-routed WebSocket hosts in 11 regions and 35+ languages.","status":"GA","models":[{"name":"falcon-2 (query param model=falcon-2)","status":"GA","notes":"~100 ms latency (docs)."},{"name":"FALCON","status":"GA","notes":"~130 ms TTFA (docs); vendor claims 55 ms model latency."},{"name":"Gen2","status":"GA","notes":"Higher-quality non-streaming model; third-party price $0.03/1K chars."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text","output":"WAV or PCM (16-bit LE), example sample_rate 24000, MONO"},"languages":"35+ (vendor)","voices":"150+ voices; enterprise voice cloning","latency":"Vendor: Falcon 55 ms model latency / 130 ms TTFA; Falcon 2 sub-100 ms TTFA.","features":["input streaming with contexts","predictive chunking and buffer controls","word timestamps (add_word_timestamps)","regional hosts","voice styles"],"pricing":{"model":"per-character","items":[{"what":"Falcon / Falcon 2","price":"$0.01 per minute","unit":"per minute (vendor headline)","notes":"Third-party trackers list $0.01 per 1K characters"},{"what":"Gen2","price":"$0.03 (third-party)","unit":"per 1K characters","notes":"Unverified"}],"est_per_minute_usd":{"low":0.009,"high":0.01,"basis":"Vendor headline 1 cent/min; 900 chars x $0.01/1K = $0.009"},"free_tier":"Third-party sources conflict ($10 monthly credit vs 100K-character trial)","source":"https://murf.ai/falcon"},"limits":["Streaming concurrency: 5 on US-East, 2 on all other regions (global router follows regional limits)","Up to 10x your concurrency in open WebSocket connections","Idle sessions closed after 3 minutes"],"regions":"global, us-east, us-west, in, ca, kr, me, jp, au, eu-central, uk, sa-east","setup":{"steps":["Get an API key from the Murf API dashboard.","Connect to wss://global.api.murf.ai/v1/speech/stream-input?api-key=...&model=falcon-2&format=PCM&sample_rate=24000.","Send a voice_config frame, then text frames with end:true on the last one.","Decode base64 'audio' fields until 'final'."],"endpoint":"wss://global.api.murf.ai/v1/speech/stream-input","auth":"api-key query parameter (invalid key closes with code 1008 after upgrade)","snippet_lang":"javascript","snippet":"// npm i ws\nimport WebSocket from \"ws\";\nimport fs from \"fs\";\n\nconst q = new URLSearchParams({ \"api-key\": process.env.MURF_API_KEY, model: \"falcon-2\",\n  sample_rate: \"24000\", channel_type: \"MONO\", format: \"PCM\" });\nconst ws = new WebSocket(`wss://global.api.murf.ai/v1/speech/stream-input?${q}`);\nconst out = fs.createWriteStream(\"out_24k_s16le.pcm\");\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({ context_id: \"t1\",\n    voice_config: { voiceId: \"YOUR_VOICE_ID\", locale: \"en-US\", style: \"Conversation\" } }));\n  ws.send(JSON.stringify({ context_id: \"t1\", text: \"Hello there. \" }));\n  ws.send(JSON.stringify({ context_id: \"t1\", text: \"Thanks for calling.\", end: true }));\n});\nws.on(\"message\", (raw) => {\n  const m = JSON.parse(raw.toString());\n  if (m.audio) out.write(Buffer.from(m.audio, \"base64\"));\n  if (m.final) ws.close();\n});\nws.on(\"close\", () => out.end());"},"warnings":[{"severity":"high","title":"Very low default concurrency outside US-East","detail":"Only 2 concurrent streams on non-US-East regions (5 on US-East). The cheap per-minute price does not help if calls queue; ask Murf for limits before launch."},{"severity":"medium","title":"API key in the URL","detail":"The documented auth puts the key in a query parameter, which can leak into proxy and access logs. Keep it server-side only."},{"severity":"medium","title":"Settings must be top-level","detail":"predictive_chunking, min_buffer_size, max_buffer_delay_in_ms and add_word_timestamps nested inside voice_config fail silently."},{"severity":"medium","title":"Latency claims differ by page","detail":"Marketing says 55 ms model latency; docs say ~130 ms (Falcon) and ~100 ms (Falcon 2) TTFA. Benchmark from your region."},{"severity":"low","title":"Pricing page did not render","detail":"murf.ai/api/pricing returned no readable prices; the 1 cent/min figure is from the Falcon marketing page."}],"best_for":"Price-driven multilingual voice agents, especially in India and other regions served by Murf's regional hosts.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"WebSocket docs","url":"https://murf.ai/api/docs/text-to-speech/web-sockets"},{"label":"Falcon","url":"https://murf.ai/falcon"}],"sources":["https://murf.ai/falcon","https://murf.ai/api/docs/text-to-speech/web-sockets","https://costbench.com/software/voice-apis/murf-api/"],"confidence":"medium","unverified":"Per-character pricing and free tier (third-party); full sample-rate list; whether model=FALCON is accepted on the socket.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":100,"languages":35,"max_session_min":null,"concurrency":5,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":true,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":10,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":true,"emotion_control":true,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Falcon 2 sub-100 ms TTFA (docs); Falcon ~130 ms, marketing says 55 ms model latency. Price: vendor headline 1 cent per minute; $10/1M is from third-party trackers. 35+ languages. Concurrency 5 on US-East, 2 in other regions. Emotion control means voice styles. EU region available."}},{"id":"speechify-api","name":"SpeechifyAI (Simba 3.2 / 3.0)","vendor":"Speechify","category":"tts","summary":"Developer API from the Speechify reader company. Cheap per-character pricing ($6-10 per 1M) and a streaming HTTP endpoint; Simba 3.2 is English-only and streaming-native. No text-input WebSocket documented.","status":"GA","models":[{"name":"simba-3.2","status":"GA (recommended for English)","notes":"Vendor: 56 ms p50 first byte (US East, 2026-09-15); Coval 106 ms and Voice Arena 123 ms medians including network (2026-09-24)."},{"name":"simba-3.0","status":"GA (default when model omitted)","notes":"English, German, Spanish, French, Italian, Portuguese."},{"name":"simba-english / simba-multilingual","status":"Retired","notes":"New workspaces get 400 model_retired."}],"transports":["HTTP chunked"],"audio":{"input":"Text and SSML (pauses, speaking rate)","output":"5 formats including mp3 (per marketing page)"},"languages":"English (simba-3.2) plus 5 more on simba-3.0","voices":"Catalog voices (e.g. geffen_32); cloning on paid plans, cloned voices on simba-3.2 need manual approval","latency":"Vendor: 56 ms p50 first byte; independent: 106 ms (Coval) and 123 ms (Voice Arena) median time to first audio.","features":["streaming HTTP output","SSML","voice cloning (paid)","batch synthesis (Scale)","docs MCP server"],"pricing":{"model":"per-character","items":[{"what":"Free","price":"$0","unit":"monthly","notes":"500K characters/month, no card; hard stop at balance"},{"what":"Starter","price":"$10/month","unit":"monthly","notes":"1.9M chars included, then $10 per 1M"},{"what":"Pro","price":"$99/month","unit":"monthly","notes":"13.5M chars included, then $8 per 1M"},{"what":"Scale","price":"$499/month","unit":"monthly","notes":"78M chars included, then $6 per 1M"}],"est_per_minute_usd":{"low":0.0054,"high":0.009,"basis":"900 chars/min at $6 (Scale) to $10 (Starter) per 1M"},"free_tier":"500K characters/month; commercial use on every plan (vendor)","source":"https://speechify.ai/pricing"},"limits":["/v1/audio/stream up to 20,000 characters; /v1/audio/speech up to 2,000","Free tier cannot top up"],"regions":"Production US East referenced for latency","setup":{"steps":["Sign up at platform.speechify.ai and create an API key.","POST https://api.speechify.ai/v1/audio/stream with input, voice_id and model.","Stream the response body to a player."],"endpoint":"https://api.speechify.ai/v1/audio/stream","auth":"Authorization: Bearer <SPEECHIFY_API_KEY>","snippet_lang":"python","snippet":"# pip install requests\nimport os, requests\n\nr = requests.post(\n    \"https://api.speechify.ai/v1/audio/stream\",\n    headers={\"Authorization\": f\"Bearer {os.environ['SPEECHIFY_API_KEY']}\",\n             \"Content-Type\": \"application/json\", \"Accept\": \"audio/mpeg\"},\n    json={\"input\": \"Streaming speech from the Speechify API.\",\n          \"voice_id\": \"geffen_32\", \"model\": \"simba-3.2\"},\n    stream=True, timeout=60)\nr.raise_for_status()\nwith open(\"out.mp3\", \"wb\") as f:\n    for chunk in r.iter_content(4096):\n        f.write(chunk)"},"warnings":[{"severity":"medium","title":"No text-input streaming","detail":"Only HTTP output streaming is documented; you must sentence-chunk LLM output and issue one request per chunk."},{"severity":"medium","title":"simba-3.2 is English only","detail":"Multilingual needs simba-3.0, which is also the silent default when you omit model."},{"severity":"medium","title":"Old model IDs fail","detail":"simba-english and simba-multilingual return 400 model_retired on new workspaces; update older sample code."},{"severity":"low","title":"Do not confuse with the consumer app","detail":"speechify.com is the reader app with separate billing; the API lives at speechify.ai / api.speechify.ai."},{"severity":"low","title":"Accept header for MP3 is assumed","detail":"The snippet sets Accept: audio/mpeg; check the API reference for the exact format selector."}],"best_for":"Low-cost English narration and agents where HTTP streaming per sentence is acceptable.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Streaming guide","url":"https://docs.speechify.ai/build/streaming-tts-guide"},{"label":"llms.txt","url":"https://docs.speechify.ai/llms.txt"}],"sources":["https://speechify.ai/pricing","https://speechify.ai/text-to-speech-api","https://docs.speechify.ai/build/streaming-tts-guide.md"],"confidence":"high","unverified":"Exact audio format parameter name and list; concurrency limits.","cat":"tts","kind":"tts","verified_at":"2026-10-10","short":"Speechify","facts":{"cat":"tts","latency_ms":56,"languages":6,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":5.26,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":false,"word_timestamps":null,"emotion_control":null,"ssml":true,"telephony_8k":null,"commercial_free_tier":true,"_notes":"56 ms p50 vendor claim (independent medians 106-123 ms). Price is the Starter plan effective rate ($10 for 1.9M chars); overage $10/1M, Scale $6/1M. simba-3.2 is English only; simba-3.0 covers 6 languages. HTTP output streaming only. Free 500K chars/month."}},{"id":"minimax-speech","name":"MiniMax Speech (international)","vendor":"MiniMax","category":"tts","summary":"High-quality multilingual TTS (40 language_boost options) with a bidirectional WebSocket that buffers LLM text and splits on punctuation. Among the most expensive per character.","status":"GA","models":[{"name":"speech-2.8-hd","status":"GA","notes":"$100/M characters."},{"name":"speech-2.8-turbo","status":"GA","notes":"$60/M characters."},{"name":"speech-2.6-hd / speech-2.6-turbo","status":"Legacy","notes":"Same prices as 2.8 equivalents."},{"name":"speech-02-hd / speech-02-turbo / speech-01-*","status":"Legacy","notes":""}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text with pause markers <#x#>","output":"mp3 (default), pcm, flac, wav, pcmu_raw, pcmu_wav, opus; 8000-44100 Hz (default 32000); mono or stereo"},"languages":"40 language_boost values plus auto","voices":"System voices; rapid voice cloning ($1.5/voice) and voice design ($3/voice); voice mixing up to 4","latency":"No millisecond figure on the bidi docs.","features":["input streaming (t2a_v2_bidi)","server-side sentence segmentation","task_cancel for barge-in","voice cloning","voice design","emotion control","voice mixing"],"pricing":{"model":"per-character","items":[{"what":"speech-2.8-hd","price":"$100","unit":"per 1M characters","notes":"Sync and async"},{"what":"speech-2.8-turbo","price":"$60","unit":"per 1M characters","notes":""},{"what":"Rapid voice cloning","price":"$1.5","unit":"per voice","notes":"Charged on first use"},{"what":"Voice design","price":"$3","unit":"per voice","notes":"Charged on first use"},{"what":"Audio subscription","price":"$5 to $999/month","unit":"monthly","notes":"100K to 20M audio points; RPM 10 to 800"}],"est_per_minute_usd":{"low":0.054,"high":0.09,"basis":"900 chars/min; turbo $60/M vs hd $100/M pay-as-you-go"},"free_tier":"Not verified","source":"https://platform.minimax.io/docs/guides/pricing-paygo.md"},"limits":["task_continue text under 10,000 characters","~120 s idle timeout; server sends no pings, client must","One synthesis session per connection","Subscription RPM: Starter 10 ... Business 800"],"regions":"International host api.minimax.io (China platform uses api.minimaxi.com with separate keys)","setup":{"steps":["Create an account on platform.minimax.io (international) and an API key.","Connect to wss://api.minimax.io/ws/v1/t2a_v2_bidi.","Wait for connected_success, send task_start (model, voice_setting, audio_setting), wait for task_started.","Send task_continue per text chunk, then task_finish; decode audio from data.audio."],"endpoint":"wss://api.minimax.io/ws/v1/t2a_v2_bidi","auth":"Authorization: Bearer <MINIMAX_API_KEY> (assumed from the standard T2A API; not shown on the bidi page excerpt)","snippet_lang":"python","snippet":"# pip install websockets\nimport asyncio, json, os, websockets\n\nURL = \"wss://api.minimax.io/ws/v1/t2a_v2_bidi\"\nHDR = {\"Authorization\": f\"Bearer {os.environ['MINIMAX_API_KEY']}\"}\n\nasync def main():\n    async with websockets.connect(URL, additional_headers=HDR) as ws:\n        await ws.recv()  # connected_success\n        await ws.send(json.dumps({\"event\": \"task_start\", \"model\": \"speech-2.8-turbo\",\n            \"voice_setting\": {\"voice_id\": \"YOUR_VOICE_ID\", \"speed\": 1},\n            \"audio_setting\": {\"format\": \"pcm\", \"sample_rate\": 24000}}))\n        await ws.recv()  # task_started\n        for t in [\"Hello there. \", \"This is streamed \", \"from an LLM.\"]:\n            await ws.send(json.dumps({\"event\": \"task_continue\", \"text\": t}))\n        await ws.send(json.dumps({\"event\": \"task_finish\"}))\n        with open(\"out_24k.pcm\", \"wb\") as f:\n            async for raw in ws:\n                m = json.loads(raw)\n                audio = (m.get(\"data\") or {}).get(\"audio\")\n                if audio:\n                    f.write(bytes.fromhex(audio))  # T2A returns hex-encoded audio\n                if m.get(\"event\") == \"task_finished\":\n                    break\n\nasyncio.run(main())"},"warnings":[{"severity":"high","title":"Expensive per character","detail":"At $60-100 per 1M characters MiniMax costs roughly 4-7x Deepgram, Inworld or Google Chirp for the same text. Model your volume first."},{"severity":"medium","title":"Short unpunctuated text waits","detail":"The bidi server synthesizes immediately only at sentence-final punctuation; short fragments without punctuation wait for a backstop window. Send task_flush at end of turn."},{"severity":"medium","title":"You must send keepalive pings","detail":"The server sends no pings and closes after ~120 s of inactivity (error 2201)."},{"severity":"medium","title":"China vs international keys","detail":"Keys from minimaxi.com (China) and minimax.io (international) are separate platforms; use the host matching your key."},{"severity":"low","title":"Audio may be hex, not base64","detail":"The standard T2A API returns hex-encoded audio; verify encoding on the bidi stream before decoding."}],"best_for":"Expressive multilingual and Chinese-language voices, cloning and voice design where cost is secondary.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Bidirectional streaming","url":"https://platform.minimax.io/docs/api-reference/speech-t2a-websocket-bidi.md"},{"label":"Pay as you go pricing","url":"https://platform.minimax.io/docs/guides/pricing-paygo.md"},{"label":"Audio subscription","url":"https://platform.minimax.io/docs/guides/pricing"}],"sources":["https://platform.minimax.io/docs/guides/pricing-paygo.md","https://platform.minimax.io/docs/guides/pricing","https://platform.minimax.io/docs/api-reference/speech-t2a-websocket-bidi.md","https://docs.livekit.io/reference/python/livekit/plugins/minimax/"],"confidence":"medium","unverified":"Auth header on bidi endpoint; hex vs base64 audio on bidi; free credits.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":null,"languages":40,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":60,"voices":null,"voice_cloning":true,"instant_cloning":true,"text_streaming_input":true,"word_timestamps":null,"emotion_control":true,"ssml":null,"telephony_8k":true,"commercial_free_tier":null,"_notes":"$60/1M is speech-2.8-turbo; hd is $100/1M. 40 language_boost values. Rapid voice cloning $1.5 per voice. Limits are RPM by subscription, not concurrency. pcmu output at 8 kHz supported."}},{"id":"neuphonic","name":"Neuphonic API","vendor":"Neuphonic","category":"tts","summary":"UK voice company focused on on-device and on-prem TTS (NeuTTS open models that run on CPU) with a hosted SSE/WebSocket API in 37 languages. API pricing is not published.","status":"GA","models":[{"name":"NeuTTS-2M / NeuTTS-2E","status":"Listed on site","notes":"~60M active params, 2x real-time on 2 CPU threads (site spec table)."},{"name":"neutts-air / neutts-nano (open weights)","status":"Open weights","notes":"See the NeuTTS open-model entry."}],"transports":["WebSocket","SSE"],"audio":{"input":"Text with inline <pause> and <phoneme> tags","output":"Not verified"},"languages":"37 (homepage playground); NeuTTS-2M spec lists 7","voices":"Instant voice cloning","latency":"Third-party listings cite a vendor sub-25 ms figure; Vapi Humanness Index measured ~763 ms for the model it tested.","features":["SSE and WebSocket streaming","voice cloning","phoneme-level pronunciation tags","on-device and on-prem deployment"],"pricing":{"model":"subscription","items":[{"what":"API","price":"Not published","unit":"","notes":"Contact sales per third-party listings"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Pricing not published"},"free_tier":"Playground available; API free tier not verified","source":"https://www.neuphonic.com/"},"limits":["Cloning samples: MP3/WAV, min 6 s, under 10 MB (third-party)"],"regions":"Not specified","setup":{"steps":["Get an API key from app.neuphonic.com.","pip install pyneuphonic (Python) or @neuphonic/neuphonic-js.","Use the SSE client for one-shot streaming or the WebSocket client for multi-chunk input."],"endpoint":"See docs.neuphonic.com","auth":"API key","snippet_lang":"python","snippet":""},"warnings":[{"severity":"medium","title":"No public price list","detail":"API pricing is quote-based; budget cannot be estimated from public pages."},{"severity":"medium","title":"Latency claims vary wildly","detail":"A vendor sub-25 ms figure circulates while an independent benchmark measured ~763 ms; test before choosing for live agents."},{"severity":"low","title":"Language counts disagree","detail":"Homepage says 37 languages; model spec lists 7 and directories list 7-10."}],"best_for":"Teams that want the same voices hosted and on-device/on-prem (CPU) for privacy.","open_source":false,"self_hostable":true,"compliance":"Not verified.","docs":[{"label":"Docs","url":"https://docs.neuphonic.com/"}],"sources":["https://www.neuphonic.com/","https://docs.neuphonic.com/","https://humannessindex.vapi.ai/providers/neuphonic","https://docs.pipecat.ai/api-reference/server/services/tts/neuphonic"],"confidence":"low","unverified":"Pricing, output formats, exact endpoints, latency.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":25,"languages":37,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":true,"open_weights":true,"price_per_1m_chars":null,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":null,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Sub-25 ms is a vendor figure cited by third parties; an independent benchmark measured ~763 ms. 37 languages per homepage, model spec lists 7. Pricing quote-based. Open weights are the NeuTTS Air/Nano models. Also SSE."}},{"id":"smallest-ai-lightning","name":"Smallest.ai Waves Lightning v3.1","vendor":"Smallest AI","category":"tts","summary":"44.1 kHz TTS with 217 voices, instant cloning, and HTTP/SSE/WebSocket transports. Default account concurrency is just 1 active TTS request.","status":"GA","models":[{"name":"lightning-v3.1","status":"GA","notes":"217 voices, 12 trained languages (20 accepted codes + auto)."},{"name":"lightning-v3.1-pro","status":"GA","notes":"Premium voice pool on dedicated capacity; same limits."},{"name":"lightning-v2 / lightning-large","status":"Deprecated","notes":"Docs say mention only for migration."}],"transports":["WebSocket","SSE","HTTP chunked"],"audio":{"input":"Text (pronunciation dictionaries on WebSocket)","output":"pcm default; 8000/16000/24000/44100 Hz"},"languages":"12 with trained voices; 20 accepted codes plus auto","voices":"217 catalog voices; instant cloning from 5-15 s","latency":"Vendor: 200 ms TTFB at 40 concurrent requests in-region. Vapi measured 420 ms median including network.","features":["input streaming with context_id continuations","word timestamps (WebSocket only)","voice cloning","pronunciation dictionaries","zero data retention (enterprise)","region-pinned hosts (India, US)"],"pricing":{"model":"per-character","items":[{"what":"Lightning v3.1","price":"Not confirmed","unit":"","notes":"Third-party: ~$14.50/1M (v3.1) and ~$19.50/1M (Pro); vendor blog: ~$0.25 per 10K chars. Agent-stack TTS layer listed at ~$0.09/min."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Official per-character rate not retrievable; third-party figures imply ~$0.013-0.023/min"},"free_tier":"$10 free credits (platform)","source":"https://smallest.ai/pricing"},"limits":["Default: 1 concurrent TTS request per account across all endpoints","Up to 5 WebSocket connections, still 1 active request","Pay-as-you-go platform concurrency listed as 20 for agents"],"regions":"India (Mumbai, api.india.smallest.ai) and USA (Oregon, api.us.smallest.ai), auto-routed","setup":{"steps":["Create an API key in the Waves dashboard.","Connect to wss://api.smallest.ai/waves/v1/tts/live with Authorization: Bearer.","Send JSON requests with text, voice_id, sample_rate, output_format, optionally context_id for continuations.","Decode base64 data.audio messages."],"endpoint":"wss://api.smallest.ai/waves/v1/tts/live","auth":"Authorization: Bearer <SMALLEST_API_KEY>","snippet_lang":"javascript","snippet":"// npm i ws  - shape follows the Lightning v3.1 model card example\nimport WebSocket from \"ws\";\nimport fs from \"fs\";\n\nconst ws = new WebSocket(\"wss://api.smallest.ai/waves/v1/tts/live\", {\n  headers: { Authorization: `Bearer ${process.env.SMALLEST_API_KEY}` },\n});\nconst out = fs.createWriteStream(\"out_24k.pcm\");\n\nws.on(\"open\", () => {\n  ws.send(JSON.stringify({ text: \"Hello from Lightning.\", voice_id: \"YOUR_VOICE_ID\",\n    sample_rate: 24000, output_format: \"pcm\" }));\n});\nws.on(\"message\", (raw) => {\n  const msg = JSON.parse(raw.toString());\n  if (msg.data?.audio) out.write(Buffer.from(msg.data.audio, \"base64\"));\n  if (msg.status === \"complete\") ws.close();\n});\nws.on(\"close\", () => out.end());"},"warnings":[{"severity":"high","title":"Concurrency of 1 by default","detail":"Only one TTS request can be processing per account at a time; a second returns 429. Vendor suggests ~4 parallel conversations at best. Get an enterprise limit before production."},{"severity":"medium","title":"Pricing is unclear","detail":"The pricing page does not list a Lightning per-character rate; third-party numbers disagree. Get a quote."},{"severity":"medium","title":"context_id cannot combine with flush","detail":"Continuations via context_id cannot be combined with flush or max_buffer_flush_ms."},{"severity":"low","title":"8 of 20 language codes reuse other voices","detail":"Only 12 languages have trained voices; the other 8 route through English or Hindi voices."}],"best_for":"Indian-language and English agents wanting cloning and 44.1 kHz audio, once concurrency is raised.","open_source":false,"self_hostable":false,"compliance":"Zero data retention on enterprise plans.","docs":[{"label":"Lightning v3.1 model card","url":"https://docs.smallest.ai/models/model-cards/text-to-speech/lightning-v-3-1.md"},{"label":"Concurrency and limits","url":"https://docs.smallest.ai/api-reference/concurrency-and-limits.md"}],"sources":["https://docs.smallest.ai/models/model-cards/text-to-speech/lightning-v-3-1.md","https://docs.smallest.ai/api-reference/concurrency-and-limits.md","https://smallest.ai/pricing","https://humannessindex.vapi.ai/models/smallestai-lightning-v31"],"confidence":"medium","unverified":"Per-character price; exact response message shape (status field name) in the snippet.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":200,"languages":12,"max_session_min":null,"concurrency":1,"free_tier":true,"free_credit_usd":10,"hipaa":null,"soc2":null,"gdpr_eu":false,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":null,"voices":217,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":true,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"200 ms TTFB at 40 concurrent in-region (vendor); Vapi measured 420 ms. Official per-character price not published; third parties say ~$14.50/1M. 12 trained languages (20 codes accepted). Default concurrency 1 per account. Hosts in India and US only."}},{"id":"playht","name":"PlayHT (Play.ai)","vendor":"PlayHT (team acqui-hired by Meta)","category":"tts","summary":"Former voice-cloning TTS API. Meta took the team in July 2025; the API went offline in July 2025 and the platform closed on December 31, 2025.","status":"Shut down","models":[{"name":"Play3.0-mini, PlayDialog","status":"Shut down","notes":"Also retired on Groq (playai-tts) in Dec 2025."}],"transports":[],"audio":{"input":"N/A","output":"N/A"},"languages":"N/A","voices":"N/A","latency":"N/A","features":[],"pricing":{"model":"subscription","items":[],"est_per_minute_usd":{"low":null,"high":null,"basis":"Service shut down"},"free_tier":"N/A","source":"https://texttolab.com/blog/play-ht-shutdown-alternatives"},"limits":[],"regions":"N/A","setup":{"steps":["Remove PlayHT integrations; migrate to another provider."],"endpoint":"N/A","auth":"N/A","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"Dead service","detail":"Multiple sources report the API went offline on July 26, 2025, earlier than the announced December 31, 2025 shutdown, and user data was deleted without export tools."},{"severity":"medium","title":"Lesson for vendor choice","detail":"Keep your own copies of voice-clone source audio and abstract the TTS provider so a vendor exit is a config change, not a rewrite."},{"severity":"medium","title":"PlayAI models also gone from Groq","detail":"Groq hosted PlayAI voices (playai-tts) until December 2025; that path is retired too and now serves Canopy Labs Orpheus instead."}],"best_for":"N/A","open_source":false,"self_hostable":false,"compliance":"N/A","docs":[],"sources":["https://texttolab.com/blog/play-ht-shutdown-alternatives","https://anyspeech.io/playht-alternatives","https://www.aiwiki.ai/wiki/playht"],"confidence":"medium","unverified":"Shutdown dates come from third-party reporting, not an official PlayHT page.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"grpc":null,"self_hostable":false,"open_weights":false,"price_per_1m_chars":null,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":null,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Service has shut down (API offline since July 2025 per third-party reports)."}},{"id":"groq-tts","name":"Groq TTS (Canopy Labs Orpheus)","vendor":"Groq","category":"tts","summary":"Groq retired its PlayAI TTS models in December 2025 and now hosts Canopy Labs' Orpheus models on its OpenAI-compatible speech endpoint. Inputs are capped at 200 characters and output is WAV only.","status":"GA","models":[{"name":"canopylabs/orpheus-v1-english","status":"GA","notes":"Vocal direction tags like [cheerful]; voices autumn, diana, hannah, austin, daniel, troy."},{"name":"canopylabs/orpheus-arabic-saudi","status":"GA","notes":"Saudi dialect; no vocal directions; terms must be accepted in console."},{"name":"playai-tts / playai-tts-arabic","status":"Deprecated / removed (Dec 2025)","notes":"Migrate to Orpheus."}],"transports":["HTTP chunked"],"audio":{"input":"Text, max 200 characters, with [directions]","output":"wav only"},"languages":"English; Arabic (Saudi)","voices":"6 English, 4 Arabic; no cloning","latency":"No numeric claim on Groq's page.","features":["vocal direction tags","OpenAI-compatible /audio/speech"],"pricing":{"model":"per-character","items":[{"what":"orpheus-v1-english","price":"$22","unit":"per 1M characters","notes":""},{"what":"orpheus-arabic-saudi","price":"$40","unit":"per 1M characters","notes":""}],"est_per_minute_usd":{"low":0.0198,"high":0.036,"basis":"900 chars/min; English vs Arabic"},"free_tier":"Groq free tier rate limits apply (not re-verified for TTS)","source":"https://console.groq.com/docs/text-to-speech/orpheus"},"limits":["200 characters max per request","WAV output only","speed parameter not supported"],"regions":"Not specified","setup":{"steps":["Create a Groq API key.","For Arabic, accept the model terms in the Groq console first.","POST to https://api.groq.com/openai/v1/audio/speech with model, voice, input (<=200 chars)."],"endpoint":"https://api.groq.com/openai/v1/audio/speech","auth":"Authorization: Bearer <GROQ_API_KEY>","snippet_lang":"python","snippet":"# pip install groq\nimport os\nfrom groq import Groq\n\nclient = Groq(api_key=os.environ[\"GROQ_API_KEY\"])\nresp = client.audio.speech.create(\n    model=\"canopylabs/orpheus-v1-english\",\n    voice=\"troy\",\n    input=\"[cheerful] Your table is ready. See you soon!\",  # max 200 chars\n    response_format=\"wav\",\n)\nresp.write_to_file(\"out.wav\")"},"warnings":[{"severity":"high","title":"200-character limit per request","detail":"Every request is capped at 200 characters, so LLM replies must be split into many calls; each adds round-trip latency and counts against rate limits."},{"severity":"high","title":"PlayAI models are gone","detail":"playai-tts and playai-tts-arabic were retired in Dec 2025; old code using them fails."},{"severity":"medium","title":"WAV only, no speed control","detail":"Only WAV output is supported and the speed parameter is ignored/unsupported."},{"severity":"medium","title":"No input streaming","detail":"Plain request/response over HTTP; no WebSocket TTS."}],"best_for":"Quick expressive English clips in apps already using Groq; not ideal for long live agent turns.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Orpheus on Groq","url":"https://console.groq.com/docs/text-to-speech/orpheus"},{"label":"Deprecations","url":"https://console.groq.com/docs/deprecations"}],"sources":["https://console.groq.com/docs/text-to-speech/orpheus","https://console.groq.com/docs/text-to-speech","https://releases.sh/release/rel_slohPX8XO2tKT7GyKNH-b-playai-tts-deprecated-orpheus-now-platform-wide"],"confidence":"high","unverified":"Exact PlayAI shutdown date (Dec 31 2025 per secondary sources).","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":22,"voices":6,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":false,"word_timestamps":null,"emotion_control":true,"ssml":null,"telephony_8k":false,"commercial_free_tier":null,"_notes":"Canopy Labs Orpheus: English ($22/1M) and Saudi Arabic ($40/1M). 6 English voices listed. 200 characters max per request, WAV output only. Vocal direction tags on English only."}},{"id":"unreal-speech","name":"Unreal Speech","vendor":"Unreal Speech","category":"tts","summary":"Budget TTS with a /stream endpoint that returns audio in about 0.3 s for up to 1,000 characters. No WebSocket input streaming.","status":"GA","models":[{"name":"API v8 (/stream, /speech, /synthesisTasks)","status":"GA","notes":"Some specs still reference api.v7 host."}],"transports":["HTTP chunked"],"audio":{"input":"Text","output":"MP3 (bitrate 16k-320k, default 192k); other formats not verified"},"languages":"English plus Chinese, Spanish, French, Hindi, Italian voices (V8 voice list)","voices":"Catalog voices by language/gender; no cloning documented","latency":"Vendor: /stream returns audio in ~0.3 s.","features":["streaming HTTP output","word timestamps on /speech","async long-form up to 500K chars"],"pricing":{"model":"subscription","items":[{"what":"Free","price":"$0","unit":"monthly","notes":"250K characters (~6 h)"},{"what":"Basic","price":"$49/month ($4.99 first 6 months)","unit":"monthly","notes":"3M characters"},{"what":"Plus","price":"$499/month","unit":"monthly","notes":"42M characters"},{"what":"Pro","price":"$1,499/month","unit":"monthly","notes":"150M characters"},{"what":"Enterprise","price":"$4,999/month","unit":"monthly","notes":"625M characters"}],"est_per_minute_usd":{"low":0.0072,"high":0.0147,"basis":"900 chars/min at plan-included rates: Enterprise $8/1M to Basic $16.3/1M (list price). Vendor itself assumes 750 chars/min."},"free_tier":"250K characters/month","source":"https://unrealspeech.com/pricing"},"limits":["/stream 1,000 chars per request","/speech 3,000 chars","/synthesisTasks 500,000 chars"],"regions":"Not specified","setup":{"steps":["Get an API key from unrealspeech.com.","POST JSON {Text, VoiceId, Bitrate} to https://api.v8.unrealspeech.com/stream.","Write the returned audio bytes as they arrive."],"endpoint":"https://api.v8.unrealspeech.com/stream","auth":"Authorization: Bearer <UNREAL_SPEECH_API_KEY>","snippet_lang":"python","snippet":"# pip install requests\nimport os, requests\n\nr = requests.post(\n    \"https://api.v8.unrealspeech.com/stream\",\n    headers={\"Authorization\": f\"Bearer {os.environ['UNREAL_SPEECH_API_KEY']}\"},\n    json={\"Text\": \"Hello from Unreal Speech.\", \"VoiceId\": \"Hannah\", \"Bitrate\": \"192k\"},\n    stream=True, timeout=30)\nr.raise_for_status()\nwith open(\"out.mp3\", \"wb\") as f:\n    for chunk in r.iter_content(4096):\n        f.write(chunk)"},"warnings":[{"severity":"medium","title":"1,000-character cap on /stream","detail":"Long LLM replies need splitting into several /stream calls."},{"severity":"medium","title":"Promotional first-month pricing","detail":"Basic shows $4.99/month for the first 6 months, then $49/month."},{"severity":"medium","title":"Host version confusion","detail":"Docs reference both api.v7 and api.v8 hosts; voice lists differ between versions. Use the host matching your docs version."},{"severity":"low","title":"No overage price published","detail":"The pricing page does not list per-1M overage rates or concurrency limits."}],"best_for":"Cheap English narration and simple bots that can tolerate per-sentence HTTP requests.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Docs (llms.txt)","url":"https://unrealspeech.readme.io/llms.txt"},{"label":"V8 docs","url":"https://docs.v8.unrealspeech.com/"}],"sources":["https://unrealspeech.com/pricing","https://unrealspeech.readme.io/llms.txt","https://docs.v8.unrealspeech.com/view/metadata/2sAYXCmKEh"],"confidence":"medium","unverified":"Auth header format, voice names on v8, output formats other than MP3, concurrency.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":300,"languages":6,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":16.33,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":false,"word_timestamps":true,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Price is the Basic plan list rate ($49 for 3M chars); promo $4.99/month for the first 6 months; Enterprise about $8/1M. /stream returns audio in ~0.3 s (vendor), 1,000 chars per request. Word timestamps on /speech only. Free 250K chars/month."}},{"id":"xai-grok-tts","name":"xAI Grok TTS","vendor":"xAI","category":"tts","summary":"Beta text-to-speech API from xAI with streaming and batch output in MP3/WAV/PCM/mu-law/A-law. Price shown as beta pricing; WebSocket text streaming is used by integrations such as Pipecat.","status":"Beta","models":[{"name":"Grok TTS (model ID not shown on the docs page fetched)","status":"Beta","notes":"OpenRouter lists 'grok-voice-tts-1.0'."}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text","output":"MP3, WAV, PCM, mu-law, A-law"},"languages":"Not verified","voices":"Built-in expressive voices; count not verified","latency":"Vapi measured 285 ms median for the streaming variant (June 2026, includes network).","features":["streaming output","WebSocket text streaming with text.clear for barge-in (per Pipecat integration)"],"pricing":{"model":"per-character","items":[{"what":"Grok TTS","price":"$4.20 (beta, per search snippet of docs) or $15 (OpenRouter and launch coverage)","unit":"per 1M characters","notes":"The docs page rendered with blank price cells when fetched"}],"est_per_minute_usd":{"low":0.0038,"high":0.0135,"basis":"900 chars/min at $4.20/1M (beta) vs $15/1M; unverified which applies"},"free_tier":"Not verified","source":"https://docs.x.ai/developers/models/text-to-speech"},"limits":["Reported: 15,000 characters per REST request; 600 RPM and 10 concurrent requests per team (secondary sources)"],"regions":"us-east-1","setup":{"steps":["Create an xAI API key.","See docs.x.ai voice/TTS reference for the REST and WebSocket endpoints."],"endpoint":"See https://docs.x.ai/developers/models/text-to-speech","auth":"Authorization: Bearer <XAI_API_KEY>","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"Beta pricing and limits may change","detail":"xAI marks TTS as beta; price and rate limits may change at GA, and sources disagree on the current rate ($4.20 vs $15 per 1M)."},{"severity":"medium","title":"Low concurrency reported","detail":"Secondary sources cite 10 concurrent requests per team, which is tight for production agents."},{"severity":"low","title":"Single US region","detail":"Cluster listed as us-east-1 only; expect higher latency from Europe/Asia."}],"best_for":"Experimenting if you already use Grok models; verify pricing first.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Models and pricing","url":"https://docs.x.ai/developers/models/text-to-speech"},{"label":"Pipecat xAI TTS","url":"https://docs.pipecat.ai/api-reference/server/services/tts/xai"}],"sources":["https://docs.x.ai/developers/models/text-to-speech","https://humannessindex.vapi.ai/models/grok-tts-streaming","https://openrouter.ai/x-ai/grok-voice-tts-1.0/apps"],"confidence":"low","unverified":"Price, model ID, WebSocket URL and message schema, voices and languages.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":10,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":false,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":null,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":true,"commercial_free_tier":null,"_notes":"Beta. Price unconfirmed: $4.20/1M (docs snippet) vs $15/1M (OpenRouter). Vapi measured 285 ms median incl. network; no vendor figure. Concurrency 10 per team from secondary sources. us-east-1 only. Text streaming per the Pipecat integration."}},{"id":"gradium","name":"Gradium TTS","vendor":"Gradium","category":"tts","summary":"Paris-based voice platform (reported as a Kyutai spin-out) with bidirectional WebSocket TTS/STT in English, French, German, Spanish and Portuguese, credit-based plans.","status":"Beta","models":[{"name":"Gradium TTS","status":"Public beta (per site banner)","notes":""}],"transports":["WebSocket"],"audio":{"input":"Text","output":"Not verified"},"languages":"English, French, German, Spanish, Portuguese (AWS Marketplace listing)","voices":"Not verified","latency":"Vendor: ~200 ms first audio; vendor-reported Coval P50 TTFA 158 ms.","features":["bidirectional WebSocket","multiplexing"],"pricing":{"model":"subscription","items":[{"what":"Free","price":"$0","unit":"monthly","notes":"45K credits (~1 h); no commercial use; TTS concurrency 2"},{"what":"XS","price":"$13/month","unit":"monthly","notes":"225K credits; concurrency 5; add-on $6.9/100K"},{"what":"S","price":"$43/month","unit":"monthly","notes":"900K credits; add-on $5.0/100K"},{"what":"M","price":"$340/month","unit":"monthly","notes":"9M credits; concurrency 10; add-on $4.0/100K"},{"what":"L","price":"$1,615/month","unit":"monthly","notes":"45M credits; concurrency 15; add-on $3.8/100K"},{"what":"TTS credit rate","price":"1 credit","unit":"per character","notes":"~45,000 characters per hour of audio"}],"est_per_minute_usd":{"low":0.032,"high":0.052,"basis":"900 chars/min; L plan ~$35.9/1M vs XS ~$57.8/1M"},"free_tier":"45K credits/month, non-commercial","source":"https://gradium.ai/pricing"},"limits":["TTS concurrency: Free 2, XS/S 5, M 10, L 15"],"regions":"Not verified (AWS Marketplace listing exists)","setup":{"steps":["Sign up at gradium.ai and create an API key.","Follow docs.gradium.ai for the WebSocket TTS protocol."],"endpoint":"See https://docs.gradium.ai","auth":"API key","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"Free tier is non-commercial","detail":"The Free plan explicitly excludes commercial use."},{"severity":"medium","title":"Expensive per character at small plans","detail":"XS works out to ~$58 per 1M characters, several times Deepgram or Inworld."},{"severity":"medium","title":"Public beta","detail":"The site banner describes the new TTS model as public beta; expect changes."}],"best_for":"European-language agents wanting a Kyutai-lineage streaming model.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Pricing","url":"https://gradium.ai/pricing"},{"label":"Docs","url":"https://docs.gradium.ai"}],"sources":["https://gradium.ai/pricing","https://aws.amazon.com/marketplace/pp/prodview-jjbvzxqmhrdiw","https://gradium.ai/content/how-to-choose-a-tts-api"],"confidence":"medium","unverified":"Endpoint URLs, audio formats, Kyutai relationship, latency.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":200,"languages":5,"max_session_min":null,"concurrency":5,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":57.8,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":false,"_notes":"Public beta. ~200 ms first audio (vendor), vendor-reported Coval P50 158 ms. Price is the XS plan effective rate ($13 for 225K credits); L plan about $36/1M. Concurrency 5 on XS (Free 2). Free plan 45K credits, non-commercial."}},{"id":"soniox-tts","name":"Soniox Text-to-Speech","vendor":"Soniox","category":"tts","summary":"Launched April 23, 2026 by the Soniox STT company: streaming TTS in 60+ languages that starts audio as text arrives, priced at about $0.70 per hour of generated speech.","status":"GA","models":[{"name":"Soniox TTS v2 (tts-rt-v2 per third-party code)","status":"GA","notes":""}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text, streamed","output":"Not verified"},"languages":"60+ (vendor launch post)","voices":"Not verified","latency":"No numeric claim verified.","features":["input streaming","character-level timestamps (per comparison page)","no storage of generated audio (vendor)"],"pricing":{"model":"per-token","items":[{"what":"Input text tokens","price":"$4.00","unit":"per 1M tokens","notes":"1 character ~ 0.3 tokens per pricing page"},{"what":"Output audio tokens","price":"$21.50","unit":"per 1M tokens","notes":""},{"what":"Approximate","price":"$0.70","unit":"per hour of generated speech","notes":"Vendor equivalent"}],"est_per_minute_usd":{"low":0.0117,"high":0.0117,"basis":"Vendor ~$0.70/hour = $0.0117/min"},"free_tier":"Not verified","source":"https://soniox.com/pricing"},"limits":["Not published for TTS"],"regions":"Not verified","setup":{"steps":["Create a Soniox API key.","Use the official SDKs (Python, Node, Web, React, React Native) or the TTS API reference."],"endpoint":"See Soniox TTS docs (third-party code references tts-rt.soniox.com)","auth":"API key","snippet_lang":"python","snippet":""},"warnings":[{"severity":"medium","title":"Token billing","detail":"Cost depends on audio tokens produced, so pacing affects price; the $0.70/hour figure is an approximation."},{"severity":"medium","title":"Young product","detail":"Launched April 2026; protocol details were not verifiable in this pass."},{"severity":"low","title":"Unverified integration details","detail":"WebSocket URL, message schema, audio formats, voices and concurrency were not confirmed from official docs here; prototype before committing."}],"best_for":"Very low-cost multilingual streaming TTS, especially if you already use Soniox STT.","open_source":false,"self_hostable":false,"compliance":"Vendor says generated audio is not stored and not used for training.","docs":[{"label":"Launch post","url":"https://soniox.com/blog/soniox-text-to-speech"},{"label":"Pricing","url":"https://soniox.com/pricing"}],"sources":["https://soniox.com/pricing","https://soniox.com/blog/soniox-text-to-speech","https://soniox.com/alternative/tts/elevenlabs"],"confidence":"medium","unverified":"Endpoints, formats, voices, concurrency.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":null,"languages":60,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":13,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":true,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Token-billed; vendor equivalent about $0.70 per hour of speech, which is about $13/1M chars at 900 chars/min. 60+ languages. Timestamps are character level. Launched April 2026; protocol details unverified. Generated audio not stored."}},{"id":"mistral-voxtral-tts","name":"Mistral Voxtral TTS","vendor":"Mistral AI","category":"tts","summary":"4B-parameter TTS released March 2026 with 20 preset voices in 9 languages and few-second voice adaptation. Hosted on Mistral's /v1/audio/speech; open weights are non-commercial (CC BY-NC 4.0).","status":"GA","models":[{"name":"voxtral-mini-tts-2603 (API) / mistralai/Voxtral-4B-TTS-2603 (weights)","status":"GA","notes":""}],"transports":["HTTP chunked"],"audio":{"input":"Text","output":"Not verified"},"languages":"9: en, fr, es, pt, it, nl, de, ar, hi","voices":"20 presets; adapts to new voices from ~3 s reference (secondary source)","latency":"Model card (self-hosted, 1 GPU): 70 ms at concurrency 1, 331 ms at 16, 552 ms at 32.","features":["streaming and batch","voice adaptation","open weights (non-commercial)"],"pricing":{"model":"per-character","items":[{"what":"Voxtral TTS","price":"$0.016","unit":"per 1K characters","notes":"Per search snippet of mistral.ai/pricing; page did not render for direct check"}],"est_per_minute_usd":{"low":0.0144,"high":0.0144,"basis":"900 chars/min x $16/1M"},"free_tier":"Not verified","source":"https://mistral.ai/pricing"},"limits":["Self-host: single GPU with >= 16 GB memory (BF16)"],"regions":"Not verified","setup":{"steps":["Create a Mistral API key.","POST to https://api.mistral.ai/v1/audio/speech with model voxtral-mini-tts-2603 (check API reference for voice parameter).","Or self-host the weights with vLLM Omni (non-commercial only)."],"endpoint":"https://api.mistral.ai/v1/audio/speech","auth":"Authorization: Bearer <MISTRAL_API_KEY>","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"Weights are non-commercial","detail":"Voxtral-4B-TTS-2603 and its reference voices are CC BY-NC 4.0; commercial products must use the paid API or a separate licence."},{"severity":"medium","title":"Latency rises fast with concurrency","detail":"Mistral's own table shows 70 ms at 1 stream but 552 ms at 32 streams per GPU."},{"severity":"low","title":"Price not directly confirmed","detail":"The $0.016/1K figure came from a search snippet of Mistral's pricing page; confirm in the console."}],"best_for":"European-language agents on Mistral; research self-hosting.","open_source":false,"self_hostable":true,"compliance":"Not verified.","docs":[{"label":"Model card","url":"https://huggingface.co/mistralai/Voxtral-4B-TTS-2603"},{"label":"Pricing","url":"https://mistral.ai/pricing"}],"sources":["https://huggingface.co/mistralai/Voxtral-4B-TTS-2603","https://mistral.ai/pricing","https://tokencost.app/models/voxtral-tts"],"confidence":"medium","unverified":"API price (snippet only), output formats, input streaming support, voice parameter name.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":70,"languages":9,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":false,"grpc":false,"self_hostable":true,"open_weights":true,"price_per_1m_chars":16,"voices":null,"voice_cloning":true,"instant_cloning":null,"text_streaming_input":null,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"70 ms is the self-hosted model card at concurrency 1 (552 ms at 32). $16/1M from a search snippet of the pricing page. Open weights are CC BY-NC 4.0 (non-commercial). Cloning means voice adaptation."}},{"id":"sarvam-bulbul","name":"Sarvam Bulbul v3","vendor":"Sarvam AI","category":"tts","summary":"Indian-language TTS (Bulbul) with a WebSocket streaming endpoint; priced in rupees per 10K characters.","status":"Beta","models":[{"name":"bulbul:v3","status":"Beta pricing","notes":"No pitch/loudness; pace 0.5-2.0; temperature; default 24 kHz."},{"name":"bulbul:v2","status":"GA","notes":""}],"transports":["WebSocket","HTTP chunked"],"audio":{"input":"Text","output":"Default 24,000 Hz; formats not verified"},"languages":"Indian languages (e.g. hi-IN) plus English; exact list not verified","voices":"Named speakers (e.g. shubh)","latency":"Not verified","features":["WebSocket streaming","flush and completion events","multiple conversions per connection"],"pricing":{"model":"per-character","items":[{"what":"Bulbul v3","price":"INR 30","unit":"per 10K characters","notes":"Beta pricing"},{"what":"Bulbul v2","price":"INR 15","unit":"per 10K characters","notes":""}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Priced in INR (about INR 2.7/min at 900 chars for v3); USD not computed to avoid an assumed FX rate"},"free_tier":"INR 1,000 free credits on every new plan","source":"https://docs.sarvam.ai/api/pricing"},"limits":["Streaming up to 2,500 characters per request","Rate limits: Starter 60 RPM, Pro 200, Business 1,000"],"regions":"India","setup":{"steps":["Create a Sarvam API key.","Use the SDK: client.text_to_speech_streaming.connect(model='bulbul:v3'), then configure(language, speaker), send text, flush."],"endpoint":"/text-to-speech/ws (Sarvam API host)","auth":"API subscription key","snippet_lang":"python","snippet":""},"warnings":[{"severity":"medium","title":"Beta pricing on v3","detail":"Bulbul v3 rates are labelled beta and may change."},{"severity":"medium","title":"Billing rounds up per request","detail":"Each request's length is rounded up, so many tiny streaming messages can cost slightly more."},{"severity":"low","title":"Send pings on long sessions","detail":"Sarvam recommends pings to keep long WebSocket connections alive."}],"best_for":"Hindi and other Indian-language voice agents.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Streaming TTS API","url":"https://docs.sarvam.ai/api-reference-docs/text-to-speech/ap-is/streaming-api"},{"label":"Pricing","url":"https://docs.sarvam.ai/api/pricing"}],"sources":["https://docs.sarvam.ai/api/pricing","https://sarvam.ai/api-pricing","https://docs.sarvam.ai/api-reference-docs/text-to-speech/ap-is/streaming-api"],"confidence":"medium","unverified":"Full language list, output formats, latency, exact host.","cat":"tts","kind":"tts","verified_at":"2026-10-10","facts":{"cat":"tts","latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":false,"webrtc":false,"websocket":true,"grpc":false,"self_hostable":false,"open_weights":false,"price_per_1m_chars":null,"voices":null,"voice_cloning":null,"instant_cloning":null,"text_streaming_input":true,"word_timestamps":null,"emotion_control":null,"ssml":null,"telephony_8k":null,"commercial_free_tier":null,"_notes":"Beta. Priced in INR: INR 30 per 10K chars (INR 3,000 per 1M) for v3, INR 15 for v2; USD not computed. Free credit INR 1,000. Indian languages plus English, list not verified. India region only. Rate limits are RPM (Starter 60)."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"kokoro","name":"Kokoro-82M","vendor":"hexgrad (community)","summary":"Tiny 82M-parameter open TTS with Apache-2.0 weights; fast enough for CPU and very fast on GPU. The default cheap self-hosted voice for many hobby and production stacks.","models":[{"name":"hexgrad/Kokoro-82M v1.0","status":"Released 2025-01-27","notes":"8 languages, 54 voices."}],"transports":["HTTP chunked"],"audio":{"input":"Text (phonemized via misaki/espeak)","output":"24 kHz float audio from the Python library; community servers expose wav/mp3/pcm"},"languages":"8 (American/British English, plus others listed in VOICES.md)","voices":"54 preset voices; no voice cloning","latency":"No official TTFB figure; generates per sentence segment.","features":["sentence-segment generator output","OpenAI-compatible community servers (e.g. Kokoro-FastAPI)","CPU capable"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/hexgrad/Kokoro-82M"},"limits":["No native text-input streaming: you feed sentences","Quality drops on unusual words (espeak fallback)"],"setup":{"steps":["pip install kokoro soundfile (and install espeak-ng for fallback/non-English).","Create a KPipeline with a lang_code.","Iterate the generator; each yield is one audio segment you can play immediately."],"endpoint":"Local (library) or your own HTTP server","auth":"None (add your own)","snippet_lang":"python","snippet":"# pip install kokoro soundfile   (+ apt install espeak-ng)\nfrom kokoro import KPipeline\nimport soundfile as sf\n\npipeline = KPipeline(lang_code=\"a\")  # 'a' = American English\ntext = \"Hello there. Kokoro yields audio sentence by sentence, so playback can start early.\"\nfor i, (graphemes, phonemes, audio) in enumerate(pipeline(text, voice=\"af_heart\")):\n    sf.write(f\"part_{i}.wav\", audio, 24000)  # stream each part to the player instead"},"hardware":"Runs on CPU; any small GPU makes it much faster than real time. 82M parameters.","license":"Apache-2.0 (weights and code)","warnings":[{"severity":"medium","title":"No cloning, fixed voices","detail":"Only the shipped voices; you cannot clone a brand voice."},{"severity":"medium","title":"Text-in streaming is your job","detail":"The pipeline splits complete text into segments. For LLM streams, buffer tokens to sentence boundaries and call it per sentence."},{"severity":"low","title":"espeak-ng dependency","detail":"Out-of-dictionary English and several non-English languages need espeak-ng installed; missing it causes silent mispronunciations or errors."},{"severity":"low","title":"Community servers vary","detail":"OpenAI-compatible wrappers are third-party projects; check their licences and maintenance."}],"best_for":"Cheapest decent-quality self-hosted English voice; edge and CPU deployments.","docs":[{"label":"Model card","url":"https://huggingface.co/hexgrad/Kokoro-82M"},{"label":"Library","url":"https://github.com/hexgrad/kokoro"}],"sources":["https://huggingface.co/hexgrad/Kokoro-82M","https://github.com/hexgrad/kokoro"],"confidence":"high","unverified":"Latency on specific hardware.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":8,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":0.082,"vram_gb":null,"cpu_ok":true,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":false,"_notes":"Sentence-segment output; no text-input streaming. 54 voices, no cloning."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"orpheus-tts","name":"Orpheus TTS (Canopy Labs)","vendor":"Canopy Labs","summary":"3B-parameter Llama-based expressive TTS with emotion tags, streaming via vLLM. Same model family Groq now hosts.","models":[{"name":"canopylabs/orpheus-3b-0.1-ft","status":"Released 2025-03","notes":"English finetune (gated, auto-approval)."},{"name":"Multilingual research release","status":"Research preview (2025-04)","notes":"7 pretrain/finetune pairs (fr, de, ko, hi, zh, es/it)."}],"transports":["HTTP chunked"],"audio":{"input":"Text with tags like <laugh>, <sigh>","output":"24 kHz 16-bit PCM chunks"},"languages":"English (prod); others research-grade","voices":"Preset voices (e.g. tara); zero-shot cloning via pretrained model","latency":"Vendor README: ~200 ms streaming latency, reducible to ~100 ms with input streaming.","features":["emotion tags","output streaming","vLLM serving","llama.cpp no-GPU option"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/canopylabs/orpheus-3b-0.1-ft"},"limits":["Needs an LLM inference stack (vLLM) for real-time"],"setup":{"steps":["pip install orpheus-speech (pulls vLLM).","Accept the gated model on Hugging Face and log in.","Call generate_speech and write chunks as they arrive."],"endpoint":"Local","auth":"Hugging Face token to download","snippet_lang":"python","snippet":"# pip install orpheus-speech   (GPU + vLLM)\nimport wave\nfrom orpheus_tts import OrpheusModel\n\nmodel = OrpheusModel(model_name=\"canopylabs/orpheus-tts-0.1-finetune-prod\")\nchunks = model.generate_speech(prompt=\"Hey there! <laugh> Long time no see.\", voice=\"tara\")\n\nwith wave.open(\"out.wav\", \"wb\") as wf:\n    wf.setnchannels(1); wf.setsampwidth(2); wf.setframerate(24000)\n    for chunk in chunks:          # PCM bytes as they are generated\n        wf.writeframes(chunk)"},"hardware":"GPU recommended (3B LLM backbone); llama.cpp path for CPU exists but is slow.","license":"Apache-2.0 per model card; built on a Llama 3.2 3B backbone, so check whether Llama licence terms also apply to your use","warnings":[{"severity":"medium","title":"Llama-derived weights","detail":"The card says Apache-2.0, but the model is built on Llama-3b. Have counsel confirm whether Meta's Llama licence obligations flow through for commercial use."},{"severity":"medium","title":"Heavy for a TTS","detail":"A 3B LLM per stream needs real GPU memory; concurrency per GPU is far lower than Kokoro-class models."},{"severity":"low","title":"Multilingual models are research previews","detail":"Non-English checkpoints are labelled research release."},{"severity":"low","title":"Hosted alternative is limited","detail":"Groq's hosted Orpheus caps requests at 200 characters and WAV only."}],"best_for":"Expressive English with laughs/sighs when you control GPUs.","docs":[{"label":"GitHub","url":"https://github.com/canopyai/Orpheus-TTS"},{"label":"Model card","url":"https://huggingface.co/canopylabs/orpheus-3b-0.1-ft"}],"sources":["https://github.com/canopyai/Orpheus-TTS","https://huggingface.co/canopylabs/orpheus-3b-0.1-ft"],"confidence":"high","unverified":"Production model name in snippet (from README) may change.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":200,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":3,"vram_gb":null,"cpu_ok":false,"commercial_ok":null,"license_short":"Apache-2.0","full_duplex":null,"streaming":true,"_notes":"About 200 ms streaming latency (about 100 ms with input streaming). Apache-2.0 per card but built on Llama 3.2 3B, so Llama licence terms may apply. Non-English checkpoints are research previews."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"chatterbox","name":"Chatterbox (Turbo, Nano, Multilingual V3)","vendor":"Resemble AI","summary":"MIT-licensed open TTS family with zero-shot cloning: Turbo (350M, English, one-step decoder, paralinguistic tags), Nano (110M, CPU) and Multilingual V3 (500M, 23+ languages). Outputs are watermarked.","models":[{"name":"ResembleAI/chatterbox-turbo","status":"Released 2025-12","notes":"350M, English, [laugh]/[cough] tags."},{"name":"ResembleAI/chatterbox-nano","status":"Released 2026-04","notes":"110M, 3x real time on 8 CPU cores (README)."},{"name":"ResembleAI/chatterbox-flash","status":"Released 2026-05","notes":"Block-diffusion variant; model card reports time-to-first-packet from 103 ms."},{"name":"Chatterbox-Multilingual V3 + single-language packs","status":"Released 2026-06","notes":"500M, 23+ languages."}],"transports":["HTTP chunked"],"audio":{"input":"Text plus optional reference audio prompt","output":"Waveform at model sample rate (model.sr)"},"languages":"Turbo/Nano: English; Multilingual V3: 23+","voices":"Zero-shot cloning from a reference clip","latency":"Hosted Resemble service claims sub-200 ms; self-hosted depends on GPU. No official streaming API in the library.","features":["zero-shot voice cloning","paralinguistic tags (Turbo/Nano)","PerTh watermark","CPU model (Nano)"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/ResembleAI/chatterbox-turbo"},"limits":["Library generates whole utterances; streaming needs community forks or your own chunking"],"setup":{"steps":["pip install chatterbox-tts.","Load ChatterboxTurboTTS on cuda (or Nano on CPU).","Generate per sentence with an optional reference clip."],"endpoint":"Local","auth":"None","snippet_lang":"python","snippet":"# pip install chatterbox-tts\nimport torchaudio as ta\nfrom chatterbox.tts_turbo import ChatterboxTurboTTS\n\nmodel = ChatterboxTurboTTS.from_pretrained(device=\"cuda\")\nwav = model.generate(\n    \"Hi there [chuckle], thanks for calling back.\",\n    audio_prompt_path=\"reference_voice.wav\",  # zero-shot clone (get consent)\n)\nta.save(\"out.wav\", wav, model.sr)"},"hardware":"Turbo: modest GPU; Nano: CPU (8 cores for 3x real time); Multilingual: GPU.","license":"MIT (Turbo, Nano, Flash, base); Dramabox model is 'other'","warnings":[{"severity":"high","title":"Zero-shot cloning needs consent","detail":"Anyone's voice can be cloned from a short clip. Get written consent and keep records; impersonation laws apply regardless of the open licence."},{"severity":"medium","title":"Built-in watermark","detail":"Every output carries Resemble's PerTh imperceptible watermark. Usually harmless, but you cannot remove it within the licence spirit."},{"severity":"medium","title":"No first-party streaming","detail":"The library returns full utterances; for agents, chunk by sentence or use a community streaming fork."},{"severity":"low","title":"Turbo and Nano are English only","detail":"Use Multilingual V3 for other languages, at higher compute."}],"best_for":"Commercial-friendly self-hosted cloning; English agents with Turbo; CPU/edge with Nano.","docs":[{"label":"GitHub","url":"https://github.com/resemble-ai/chatterbox"},{"label":"Turbo model card","url":"https://huggingface.co/ResembleAI/chatterbox-turbo"}],"sources":["https://github.com/resemble-ai/chatterbox","https://huggingface.co/ResembleAI/chatterbox-turbo","https://huggingface.co/ResembleAI/chatterbox-nano","https://huggingface.co/ResembleAI/chatterbox-flash"],"confidence":"high","unverified":"Self-hosted latency numbers.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":103,"languages":23,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":0.35,"vram_gb":null,"cpu_ok":true,"commercial_ok":true,"license_short":"MIT","full_duplex":null,"streaming":false,"_notes":"103 ms is time to first packet for the Flash variant per its card. 23+ languages on Multilingual V3 (Turbo and Nano English only). Params: Turbo 350M (Nano 110M runs on CPU, Multilingual 500M). No first-party streaming in the library. Outputs watermarked."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"sesame-csm","name":"Sesame CSM-1B","vendor":"Sesame","summary":"1B conversational speech model that conditions on prior dialogue context. Apache-2.0, in Hugging Face Transformers since 4.52.1. Not a low-latency streaming engine out of the box.","models":[{"name":"sesame/csm-1b","status":"Released 2025-03 (gated, auto-approval)","notes":"Only public Sesame checkpoint."}],"transports":[],"audio":{"input":"Text plus optional conversation context (text + audio segments)","output":"24 kHz waveform (generator.sample_rate)"},"languages":"English (other languages weak, per README)","voices":"Speaker IDs and context-based voice continuation","latency":"No streaming or latency claim in the README.","features":["context-conditioned conversational prosody","Transformers integration"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/sesame/csm-1b"},"limits":["No official streaming","English focus"],"setup":{"steps":["git clone https://github.com/SesameAILabs/csm and install requirements (CUDA GPU).","Accept the gated model on Hugging Face.","Call load_csm_1b and generate."],"endpoint":"Local","auth":"Hugging Face token","snippet_lang":"python","snippet":"# from the SesameAILabs/csm repo (run inside the cloned repo)\nimport torchaudio\nfrom generator import load_csm_1b\n\ngenerator = load_csm_1b(device=\"cuda\")\naudio = generator.generate(text=\"Hello from Sesame.\", speaker=0, context=[],\n                           max_audio_length_ms=10_000)\ntorchaudio.save(\"audio.wav\", audio.unsqueeze(0).cpu(), generator.sample_rate)"},"hardware":"CUDA GPU recommended.","license":"Apache-2.0","warnings":[{"severity":"medium","title":"The open model is not the demo","detail":"The public 1B checkpoint is a base generation model; it does not reproduce Sesame's hosted conversational demo out of the box."},{"severity":"medium","title":"Not built for streaming agents","detail":"No official streaming path; latency is high compared with purpose-built realtime TTS."},{"severity":"low","title":"Gated download","detail":"Requires accepting terms on Hugging Face before download."}],"best_for":"Research into context-aware conversational speech.","docs":[{"label":"GitHub","url":"https://github.com/SesameAILabs/csm"},{"label":"Model card","url":"https://huggingface.co/sesame/csm-1b"}],"sources":["https://github.com/SesameAILabs/csm","https://huggingface.co/sesame/csm-1b"],"confidence":"medium","unverified":"Hardware minimums.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":1,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":false,"_notes":"No official streaming path. Gated download."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"f5-tts","name":"F5-TTS","vendor":"SWivid (academic)","summary":"Popular flow-matching zero-shot cloning model. Code is MIT but the pretrained weights are CC-BY-NC because of the Emilia training data.","models":[{"name":"F5TTS_v1_Base","status":"Released 2025-03-12","notes":""}],"transports":[],"audio":{"input":"Text plus reference audio and its transcript","output":"24 kHz waveform"},"languages":"English and Chinese base; community finetunes for others","voices":"Zero-shot cloning","latency":"Non-autoregressive; no first-party streaming server.","features":["zero-shot cloning","Docker image","CLI and Gradio app"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"Weights free for non-commercial use only"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/SWivid/F5-TTS"},"limits":["Non-commercial weights"],"setup":{"steps":["pip install f5-tts (GPU).","Run f5-tts_infer-cli --model F5TTS_v1_Base with a reference clip."],"endpoint":"Local","auth":"None","snippet_lang":"python","snippet":""},"hardware":"GPU","license":"Code MIT; pretrained weights CC-BY-NC-4.0","warnings":[{"severity":"high","title":"Weights are non-commercial","detail":"The official checkpoints are CC-BY-NC due to Emilia training data. Commercial products need your own retrained weights."},{"severity":"medium","title":"Not a streaming engine","detail":"Generates whole utterances; poor fit for live agents."},{"severity":"medium","title":"Cloning consent","detail":"Zero-shot cloning of real people needs consent."}],"best_for":"Research and non-commercial cloning experiments.","docs":[{"label":"GitHub","url":"https://github.com/SWivid/F5-TTS"}],"sources":["https://github.com/SWivid/F5-TTS","https://huggingface.co/SWivid/F5-TTS"],"confidence":"high","unverified":"","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":null,"commercial_ok":false,"license_short":"CC-BY-NC-4.0","full_duplex":null,"streaming":false,"_notes":"Code MIT, pretrained weights CC-BY-NC-4.0. English and Chinese base."}},{"category":"open-model","status":"Deprecated","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"xtts-v2","name":"Coqui XTTS-v2","vendor":"Coqui (defunct); maintained fork by Idiap","summary":"Once the default open cloning TTS (17 languages, <200 ms streaming). Coqui the company is gone; the code lives on in the Idiap fork (coqui-tts on PyPI), but the weights remain under the non-commercial Coqui Public Model License.","models":[{"name":"tts_models/multilingual/multi-dataset/xtts_v2","status":"Unmaintained weights (2023)","notes":"17 languages."}],"transports":["HTTP chunked"],"audio":{"input":"Text plus 6 s+ reference audio","output":"24 kHz"},"languages":"17","voices":"Zero-shot cloning","latency":"Project claim: can stream with <200 ms latency (GPU).","features":["streaming inference","cloning","fine-tuning recipes"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"Non-commercial licence; commercial licences can no longer be bought from Coqui"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/coqui/XTTS-v2"},"limits":["CPML non-commercial"],"setup":{"steps":["pip install coqui-tts (Idiap fork).","Load TTS('tts_models/multilingual/multi-dataset/xtts_v2') on GPU."],"endpoint":"Local","auth":"None","snippet_lang":"python","snippet":"# pip install coqui-tts   (maintained Idiap fork)\nfrom TTS.api import TTS\n\ntts = TTS(\"tts_models/multilingual/multi-dataset/xtts_v2\").to(\"cuda\")\ntts.tts_to_file(text=\"Hello from XTTS.\", speaker_wav=\"reference.wav\",\n                language=\"en\", file_path=\"out.wav\")"},"hardware":"GPU for real-time","license":"Coqui Public Model License (non-commercial) for weights; code MPL-2.0 in the fork (not re-verified)","warnings":[{"severity":"high","title":"No way to license commercially","detail":"Weights are under the non-commercial CPML and Coqui no longer exists to sell a commercial licence."},{"severity":"medium","title":"Weights unmaintained since 2023","detail":"Quality and robustness lag 2026 models such as Chatterbox, Qwen3-TTS or VoxCPM2."},{"severity":"low","title":"Use the fork","detail":"The original coqui-ai/TTS repo is unmaintained; install coqui-tts from the Idiap fork."}],"best_for":"Legacy/non-commercial projects already built on it.","docs":[{"label":"Idiap fork","url":"https://github.com/idiap/coqui-ai-TTS"},{"label":"Model card","url":"https://huggingface.co/coqui/XTTS-v2"}],"sources":["https://github.com/idiap/coqui-ai-TTS","https://huggingface.co/coqui/XTTS-v2"],"confidence":"high","unverified":"Fork code licence.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":200,"languages":17,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":false,"commercial_ok":false,"license_short":"Custom","full_duplex":null,"streaming":true,"_notes":"Project claims under 200 ms streaming on GPU. Weights under the non-commercial Coqui Public Model License with no way to buy a commercial licence. Unmaintained since 2023."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"kyutai-tts","name":"Kyutai TTS (Delayed Streams Modeling)","vendor":"Kyutai","summary":"Streaming-native open TTS that starts speaking before the full text is known, served by a production Rust WebSocket server (the engine behind Kyutai's Unmute). CC-BY-4.0 weights.","models":[{"name":"kyutai/tts-1.6b-en_fr","status":"Released 2025-06","notes":"English and French."},{"name":"kyutai/tts-0.75b-en-public","status":"Released 2025-07","notes":"English."}],"transports":["WebSocket"],"audio":{"input":"Streamed text","output":"Streamed audio (Mimi codec)"},"languages":"English, French","voices":"Voice embeddings from the kyutai/tts-voices repository (each voice has its own licence)","latency":"No exact TTS TTFB on the README; the same Rust server is used for Unmute.","features":["true text-input streaming","Rust WebSocket server (moshi-server)","batching"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/kyutai/tts-1.6b-en_fr"},"limits":["English/French only"],"setup":{"steps":["Clone kyutai-labs/delayed-streams-modeling.","For production install the Rust server (moshi-server) and run: moshi-server worker --config configs/config-tts.toml.","Stream text with scripts/tts_rust_server.py or your own WebSocket client."],"endpoint":"Your moshi-server WebSocket","auth":"Your own","snippet_lang":"python","snippet":"# Quick local test with the PyTorch streaming script from the repo:\n#   echo \"Hey, how are you?\" | python scripts/tts_pytorch_streaming.py audio_output.wav\n# Production: Rust server over WebSocket\n#   moshi-server worker --config configs/config-tts.toml\n#   echo \"Hey, how are you?\" | python scripts/tts_rust_server.py - -\nimport subprocess\nsubprocess.run('echo \"Hello from Kyutai.\" | python scripts/tts_pytorch_streaming.py out.wav',\n               shell=True, check=True)"},"hardware":"NVIDIA GPU for the Rust server (L40S class used for Unmute)","license":"Weights CC-BY-4.0 (attribution required); voices have per-voice licences","warnings":[{"severity":"medium","title":"Attribution required","detail":"CC-BY-4.0 allows commercial use but requires attribution; check each voice's separate licence in kyutai/tts-voices."},{"severity":"medium","title":"Ops complexity","detail":"Production use means running the Rust moshi-server with CUDA; not a pip-install-and-go experience."},{"severity":"low","title":"Two languages only","detail":"English and French."}],"best_for":"Self-hosted agents that need genuine text-in streaming without sentence buffering.","docs":[{"label":"GitHub","url":"https://github.com/kyutai-labs/delayed-streams-modeling"},{"label":"Model","url":"https://huggingface.co/kyutai/tts-1.6b-en_fr"}],"sources":["https://github.com/kyutai-labs/delayed-streams-modeling","https://huggingface.co/kyutai/tts-1.6b-en_fr"],"confidence":"medium","unverified":"Latency and per-GPU concurrency for TTS specifically.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":2,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":1.6,"vram_gb":null,"cpu_ok":false,"commercial_ok":true,"license_short":"CC-BY-4.0","full_duplex":null,"streaming":true,"_notes":"True text-input streaming. Values for tts-1.6b-en_fr (also 0.75B English). Each voice has its own licence. Rust server needs an NVIDIA GPU."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"kyutai-pocket-tts","name":"Kyutai Pocket TTS","vendor":"Kyutai","summary":"Small CPU-only TTS with voice cloning: ~200 ms to first audio chunk and ~6x real time on a MacBook Air M4 using 2 cores. Six languages.","models":[{"name":"kyutai/pocket-tts","status":"Gated on Hugging Face; updated 2026-10-01","notes":"Also pocket-tts-without-voice-cloning."}],"transports":["HTTP chunked"],"audio":{"input":"Text plus voice prompt","output":"PCM tensor at tts_model.sample_rate"},"languages":"en, fr, de, pt, it, es (model card)","voices":"Pre-made voices and cloning from a local wav","latency":"README: ~200 ms to first audio chunk; ~6x real time on M4 CPU with 2 cores.","features":["CPU only","audio streaming","voice cloning","built-in 'serve' command"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/kyutai/pocket-tts"},"limits":["Gated model access"],"setup":{"steps":["pip install pocket-tts (on Linux use the CPU PyTorch index).","Accept the gated model on Hugging Face.","Use the Python API or 'pocket-tts serve'."],"endpoint":"Local","auth":"Hugging Face token","snippet_lang":"python","snippet":"# pip install pocket-tts\nfrom pocket_tts import TTSModel\nimport scipy.io.wavfile\n\ntts_model = TTSModel.load_model()\nvoice_state = tts_model.get_state_for_audio_prompt(\"alba\")  # or a local wav to clone\naudio = tts_model.generate_audio(voice_state, \"Hello world, this is a test.\")\nscipy.io.wavfile.write(\"output.wav\", tts_model.sample_rate, audio.numpy())"},"hardware":"CPU (2 cores); no GPU needed","license":"CC-BY-4.0 (model card)","warnings":[{"severity":"medium","title":"Gated download","detail":"Access to kyutai/pocket-tts requires approval on Hugging Face."},{"severity":"medium","title":"Cloning consent","detail":"Voice cloning from any wav needs speaker consent."},{"severity":"low","title":"Load voices once","detail":"load_model and get_state_for_audio_prompt are slow; cache them in memory."}],"best_for":"On-device or cheap CPU servers needing cloning in European languages.","docs":[{"label":"GitHub","url":"https://github.com/kyutai-labs/pocket-tts"}],"sources":["https://github.com/kyutai-labs/pocket-tts","https://huggingface.co/kyutai/pocket-tts"],"confidence":"medium","unverified":"Code licence.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":200,"languages":6,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":true,"commercial_ok":true,"license_short":"CC-BY-4.0","full_duplex":null,"streaming":true,"_notes":"200 ms to first chunk on a MacBook Air M4 CPU (2 cores), no GPU needed. Gated download."}},{"category":"open-model","status":"Beta","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"vibevoice","name":"Microsoft VibeVoice (Realtime-0.5B)","vendor":"Microsoft","summary":"MIT-licensed research TTS. The long-form VibeVoice-TTS code was pulled in Sept 2025 over misuse; VibeVoice-Realtime-0.5B (Dec 2025) supports streaming text input with ~200 ms first audible latency, English-focused and single-speaker.","models":[{"name":"microsoft/VibeVoice-Realtime-0.5B","status":"Released 2025-12-03","notes":"Streaming text input; embedded voice prompts only."},{"name":"microsoft/VibeVoice-1.5B","status":"Code disabled (2025-09-05)","notes":"Weights still on HF but TTS code removed from the repo."}],"transports":["WebSocket"],"audio":{"input":"Streamed text","output":"Streamed audio"},"languages":"English (9 experimental languages: de, fr, it, ja, ko, nl, pl, pt, es)","voices":"Embedded preset speakers only; custom voices by request to Microsoft","latency":"README: ~200 ms first audible latency (hardware dependent).","features":["streaming text input","long-form generation","demo web server"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B"},"limits":["Single speaker","No user voice cloning"],"setup":{"steps":["git clone https://github.com/microsoft/VibeVoice && pip install -e .[streamingtts]","Run python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B","Use an NVIDIA Deep Learning Container for CUDA."],"endpoint":"Local demo server","auth":"None","snippet_lang":"python","snippet":"# Shell (from the VibeVoice README):\n#   git clone https://github.com/microsoft/VibeVoice.git && cd VibeVoice\n#   pip install -e .[streamingtts]\n#   python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B\n# File-based test:\nimport subprocess\nsubprocess.run([\"python\", \"demo/realtime_model_inference_from_file.py\",\n                \"--model_path\", \"microsoft/VibeVoice-Realtime-0.5B\",\n                \"--txt_path\", \"demo/text_examples/1p_vibevoice.txt\",\n                \"--speaker_name\", \"Carter\"], check=True)"},"hardware":"NVIDIA GPU (CUDA)","license":"MIT; Microsoft frames it as a research framework","warnings":[{"severity":"high","title":"Research release with a misuse history","detail":"Microsoft removed the main VibeVoice-TTS code in Sept 2025 after misuse. Treat as research; it could change or be withdrawn again."},{"severity":"medium","title":"No custom voices","detail":"Voice prompts are embedded to limit deepfakes; you cannot clone your own voice."},{"severity":"low","title":"English only in practice","detail":"Other languages are explicitly experimental."}],"best_for":"Experimenting with LLM-token-to-speech streaming on your own GPU.","docs":[{"label":"Realtime model doc","url":"https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-realtime-0.5b.md"}],"sources":["https://github.com/microsoft/VibeVoice","https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B"],"confidence":"high","unverified":"Transport of the demo server (described as real-time service; WebSocket assumed).","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":200,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_weights":true,"params_b":0.5,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"MIT","full_duplex":null,"streaming":true,"_notes":"VibeVoice-Realtime-0.5B; 9 more languages are experimental. MIT but framed by Microsoft as a research release (long-form code was withdrawn after misuse). No voice cloning."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"dia","name":"Nari Labs Dia / Dia2","vendor":"Nari Labs","summary":"Apache-2.0 dialogue TTS that writes two-speaker conversations with non-verbal cues. Dia2 (1B/2B, Nov 2025) adds streaming; English only, up to 2 minutes.","models":[{"name":"nari-labs/Dia2-2B, Dia2-1B","status":"Released 2025-11/12","notes":"Streaming dialogue TTS."},{"name":"nari-labs/Dia-1.6B-0626","status":"Older","notes":""}],"transports":["HTTP chunked"],"audio":{"input":"Script with [S1]/[S2] speaker tags","output":"Waveform (Mimi codec for Dia2)"},"languages":"English","voices":"Speaker conditioning/cloning from prompt audio","latency":"No numeric claim; Dia2 server described as real streaming.","features":["multi-speaker dialogue","non-verbal sounds","streaming server (Dia2)"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/nari-labs/Dia2-2B"},"limits":["Up to 2 minutes of generation","English only"],"setup":{"steps":["Clone github.com/nari-labs/dia2 and follow the README (GPU).","Run the Dia2 TTS server for streaming."],"endpoint":"Local","auth":"None","snippet_lang":"python","snippet":""},"hardware":"GPU","license":"Apache-2.0 (third-party assets such as Kyutai Mimi keep their own licences)","warnings":[{"severity":"medium","title":"2-minute cap","detail":"Dia2 generates up to 2 minutes; long content must be segmented."},{"severity":"medium","title":"Identity misuse","detail":"Nari Labs explicitly prohibits producing audio resembling real people without permission."},{"severity":"low","title":"Dialogue-first","detail":"Optimised for two-speaker scripts, not single-voice agent replies."}],"best_for":"Podcast-style two-speaker dialogue generation.","docs":[{"label":"Dia2 GitHub","url":"https://github.com/nari-labs/dia2"}],"sources":["https://github.com/nari-labs/dia2","https://huggingface.co/nari-labs/Dia2-2B"],"confidence":"medium","unverified":"Latency and hardware minimums.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":2,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":true,"_notes":"Dia2-2B (also 1B). Two-speaker dialogue, up to 2 minutes of generation."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"qwen3-tts","name":"Qwen3-TTS","vendor":"Alibaba Qwen","summary":"Apache-2.0 open TTS (0.6B and 1.7B) released January 2026 with dual-track streaming that can emit audio after a single input character, 10 languages, instruction control, voice design and cloning.","models":[{"name":"Qwen/Qwen3-TTS-12Hz-1.7B-Base / -CustomVoice / 0.6B variants","status":"Released 2026-01-21","notes":""}],"transports":["HTTP chunked"],"audio":{"input":"Text plus optional instruction","output":"Waveform (sample rate returned by generate)"},"languages":"10: zh, en, ja, ko, de, fr, ru, pt, es, it","voices":"Preset speakers, voice design, voice cloning","latency":"README: end-to-end synthesis latency as low as 97 ms; first packet after one character.","features":["streaming generation","instruction-controlled emotion and pace","voice design","voice cloning","vLLM-Omni support (offline at launch)"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base"},"limits":["vLLM-Omni online serving was 'coming later' at launch"],"setup":{"steps":["pip install -U qwen-tts (flash-attn recommended).","Load a model from Hugging Face.","Call generate_custom_voice / voice design / clone functions."],"endpoint":"Local","auth":"None","snippet_lang":"python","snippet":"# pip install -U qwen-tts soundfile\nimport torch, soundfile as sf\nfrom qwen_tts import Qwen3TTSModel\n\nmodel = Qwen3TTSModel.from_pretrained(\"Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice\",\n    device_map=\"cuda:0\", dtype=torch.bfloat16, attn_implementation=\"flash_attention_2\")\nwavs, sr = model.generate_custom_voice(\n    text=\"Your appointment is confirmed for Tuesday at ten.\",\n    language=\"English\", speaker=\"Vivian\", instruct=\"Calm and friendly.\")\nsf.write(\"out.wav\", wavs[0], sr)"},"hardware":"NVIDIA GPU (bf16, flash-attention recommended)","license":"Apache-2.0","warnings":[{"severity":"medium","title":"Streaming needs the right serving path","detail":"The 97 ms figure is from the vendor; the simple Python API returns whole utterances. Real streaming depends on the serving stack you choose."},{"severity":"medium","title":"Cloning consent","detail":"Voice cloning of real people requires consent."},{"severity":"low","title":"flash-attn build pain","detail":"flash-attn often needs compiling; set MAX_JOBS to avoid running out of RAM."}],"best_for":"Commercial-friendly self-hosted multilingual (especially Chinese/English) TTS with instructions.","docs":[{"label":"GitHub","url":"https://github.com/QwenLM/Qwen3-TTS"}],"sources":["https://github.com/QwenLM/Qwen3-TTS","https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base"],"confidence":"high","unverified":"Independent latency.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":97,"languages":10,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":1.7,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":true,"_notes":"97 ms is the vendor end-to-end figure; the simple Python API returns whole utterances. Also 0.6B variants."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"voxcpm2","name":"VoxCPM2","vendor":"OpenBMB","summary":"2B tokenizer-free diffusion-autoregressive TTS with 30 languages, 48 kHz output, voice design and cloning, Apache-2.0 and explicitly commercial-ready.","models":[{"name":"openbmb/VoxCPM2","status":"Released 2026-04; updated 2026-08","notes":""}],"transports":["HTTP chunked"],"audio":{"input":"Text, optional voice description or reference audio","output":"48 kHz"},"languages":"30","voices":"Voice design from text; controllable and 'ultimate' cloning","latency":"Model card: real-time streaming with RTF ~0.3 on RTX 4090, ~0.13 with Nano-vLLM.","features":["streaming","voice design","voice cloning","multilingual without language tags"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/openbmb/VoxCPM2"},"limits":[],"setup":{"steps":["pip install voxcpm soundfile.","Load VoxCPM.from_pretrained('openbmb/VoxCPM2').","Call generate."],"endpoint":"Local","auth":"None","snippet_lang":"python","snippet":"# pip install voxcpm soundfile\nfrom voxcpm import VoxCPM\nimport soundfile as sf\n\nmodel = VoxCPM.from_pretrained(\"openbmb/VoxCPM2\", load_denoiser=False)\nwav = model.generate(text=\"VoxCPM2 brings multilingual support and voice design.\")\nsf.write(\"output.wav\", wav, model.tts_model.sample_rate)"},"hardware":"NVIDIA GPU (RTX 4090 class for the quoted RTF)","license":"Apache-2.0","warnings":[{"severity":"medium","title":"RTF is not TTFB","detail":"The card quotes real-time factor, not time to first audio; measure first-chunk latency yourself."},{"severity":"medium","title":"Cloning consent","detail":"High-fidelity cloning from a short clip needs speaker consent."},{"severity":"low","title":"2B params","detail":"Heavier than Kokoro/Chatterbox-class models; plan GPU capacity per stream."}],"best_for":"Commercial self-hosted multilingual voices with design and cloning.","docs":[{"label":"Model card","url":"https://huggingface.co/openbmb/VoxCPM2"}],"sources":["https://huggingface.co/openbmb/VoxCPM2"],"confidence":"medium","unverified":"TTFB; generate() argument names beyond the card's example.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":30,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":2,"vram_gb":null,"cpu_ok":null,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":true,"_notes":"Only real-time factor published (about 0.3 on RTX 4090), no time to first audio. 48 kHz output."}},{"category":"open-model","status":"GA","open_source":true,"self_hostable":true,"regions":"Wherever you deploy it","compliance":"Your own deployment; no vendor data processing","id":"neutts","name":"NeuTTS Air / Nano","vendor":"Neuphonic","summary":"Small on-device TTS models with instant cloning that run on CPU (GGUF via llama.cpp). Air is Apache-2.0; Nano uses a custom licence. Outputs are watermarked.","models":[{"name":"neuphonic/neutts-air (+ q4/q8 GGUF)","status":"Released 2025-09","notes":"English."},{"name":"neuphonic/neutts-nano (+ per-language variants)","status":"Released 2025-11; language packs 2026-08","notes":"English, Spanish, French, German variants."}],"transports":[],"audio":{"input":"Text plus reference audio for cloning","output":"Waveform via NeuCodec"},"languages":"Air: English; Nano: English plus es/fr/de variants","voices":"Instant cloning","latency":"Built for real-time or faster on laptop CPUs (model card); no TTFB figure.","features":["CPU and GGUF deployment","instant cloning","PerTh watermark"],"pricing":{"model":"free","items":[{"what":"Weights","price":"$0","unit":"","notes":"You pay for your own compute"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted: cost is your GPU/CPU time, not per character"},"free_tier":"Open weights","source":"https://huggingface.co/neuphonic/neutts-air"},"limits":["Gated (auto) downloads"],"setup":{"steps":["Install the neutts package from Neuphonic's GitHub and optionally llama-cpp-python for GGUF.","Load NeuTTS with backbone_repo='neuphonic/neutts-air-q4-gguf' on CPU.","Encode a reference clip and generate."],"endpoint":"Local","auth":"Hugging Face token","snippet_lang":"python","snippet":""},"hardware":"CPU (laptop class); GPU optional","license":"Air: Apache-2.0; Nano: 'other' (custom Neuphonic licence, check terms)","warnings":[{"severity":"medium","title":"Nano is not Apache","detail":"neutts-nano is listed with a custom 'other' licence; read it before commercial use."},{"severity":"low","title":"Watermarked output","detail":"NeuTTS Air embeds the Perth watermark in every file."},{"severity":"low","title":"English-centric","detail":"Air is English only; other languages need specific Nano variants."}],"best_for":"Private on-device voice with cloning on CPUs and edge hardware.","docs":[{"label":"NeuTTS Air card","url":"https://huggingface.co/neuphonic/neutts-air"},{"label":"NeuTTS Nano card","url":"https://huggingface.co/neuphonic/neutts-nano"}],"sources":["https://huggingface.co/neuphonic/neutts-air","https://huggingface.co/neuphonic/neutts-nano"],"confidence":"medium","unverified":"Package name and API; Nano licence text.","cat":"open","kind":"tts","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":1,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_weights":true,"params_b":null,"vram_gb":null,"cpu_ok":true,"commercial_ok":true,"license_short":"Apache-2.0","full_duplex":null,"streaming":null,"_notes":"Values for NeuTTS Air (English). Nano has es/fr/de variants under a custom licence. Outputs watermarked."}},{"id":"livekit","name":"LiveKit Agents + LiveKit Cloud","vendor":"LiveKit","category":"framework","summary":"Open-source (Apache-2.0) Python/Node framework for realtime voice agents on top of LiveKit's WebRTC SFU, with a managed cloud that hosts agents, provides telephony and a pay-per-use inference gateway for STT, LLM and TTS.","status":"GA","models":[{"name":"Build plan","status":"GA","notes":"$0/mo, no card: 1,000 agent session minutes, 5 concurrent agent sessions, 1 deployment, $2.50 inference credit (~50 min), 1 free US local number with 50 inbound minutes."},{"name":"Ship plan","status":"GA","notes":"$50/mo: 5,000 agent session minutes then $0.01/min, 20 concurrent sessions, $5 inference credit."},{"name":"Scale plan","status":"GA","notes":"$500/mo: 50,000 agent session minutes then $0.01/min, inference concurrency 50, $50 inference credit and discounted model prices. Pricing page lists 'up to 600' concurrent sessions but the table layout is ambiguous."},{"name":"Enterprise","status":"GA","notes":"Custom."},{"name":"LiveKit Inference (gateway)","status":"GA","notes":"One API key for many STT/LLM/TTS models, billed per minute on the LiveKit invoice (e.g. Deepgram Nova-3 $0.0048/min, AssemblyAI Universal-Streaming $0.0025/min, Cartesia Sonic 3 $0.03/min, Deepgram Aura-2 $0.018/min on Build/Ship). Some TTS models are listed as free."},{"name":"Self-hosted LiveKit server + Agents","status":"GA","notes":"Open source; you run the SFU, TURN, agent workers and pay providers directly."}],"transports":["WebRTC","SIP","PSTN"],"audio":{"input":"Opus over WebRTC; SIP/PSTN calls bridged in","output":"Opus over WebRTC; telephony codecs on SIP legs"},"languages":"Depends on the chosen STT/TTS/LLM; LiveKit Inference restricts to EU-hosted models for EU agents.","latency":"No single vendor number verified; depends on models and region.","features":["turn detection (transformer model + inference.TurnDetector)","noise cancellation (ai-coustics, BVC)","telephony (LiveKit Phone Numbers, third-party SIP, Twilio connector)","multi-provider plug-ins","recording","observability/traces","MCP tools","realtime speech-to-speech models","agent hosting (lk agent create)"],"pricing":{"model":"usage","items":[{"what":"Agent session minutes","price":"$0.01","unit":"per minute","notes":"After included minutes (1,000 Build, 5,000 Ship, 50,000 Scale). Hosting/orchestration only; excludes model costs."},{"what":"Agent session recordings","price":"$0.005","unit":"per minute","notes":"After included minutes."},{"what":"WebRTC participant minutes","price":"$0.0005 (Ship) / $0.0004 (Scale)","unit":"per minute","notes":"5,000 / 150,000 / 1.5M included on Build / Ship / Scale."},{"what":"Third-party SIP minutes","price":"$0.004 (Ship) / $0.003 (Scale)","unit":"per minute","notes":"Your own SIP trunk, carrier charges billed by your carrier."},{"what":"US local inbound (LiveKit number)","price":"$0.01","unit":"per minute","notes":"Plus $1.00/month per extra number; toll-free $2.00/month and $0.02/min inbound."},{"what":"Downstream data","price":"$0.12 (Ship) / $0.10 (Scale)","unit":"per GB","notes":"After 50GB/250GB/3TB."},{"what":"Plans","price":"$0 / $50 / $500 per month","unit":"per month","notes":"Build / Ship / Scale."},{"what":"Inference STT/LLM/TTS","price":"e.g. STT $0.0025-$0.0048, TTS $0.018-$0.03, LLM GPT-4o mini $0.0006","unit":"per minute","notes":"Optional; or bring your own provider keys via plugins."}],"est_per_minute_usd":{"low":0.03,"high":0.08,"basis":"LiveKit's own calculator default is $0.0479/min (agent session $0.01 + telephony $0.01 + STT $0.0075 + LLM $0.0014 + TTS $0.009 + observability $0.01). Low end: WebRTC only, cheap STT/TTS, small LLM. High end: phone calls, premium TTS (Cartesia/ElevenLabs class) and a larger LLM. Speech-to-speech realtime models cost much more and are priced in the model segment."},"free_tier":"Build plan: 1,000 agent session minutes/month, $2.50 inference credit, 1 US number with 50 inbound minutes, no credit card.","source":"https://livekit.com/pricing"},"limits":["Concurrent agent sessions: 5 Build, 20 Ship, up to 600 Scale (page ambiguous), custom Enterprise","Inference concurrency: 5 / 20 / 50","Agent deployments: 1 / 2 / 4","WebRTC concurrent connections: 100 / 1,000 / 5,000"],"regions":"Global SFU; project data region US (default) or EU (Frankfurt), fixed at creation; agents deployable to regions such as eu-central.","setup":{"steps":["Install the LiveKit CLI and run 'lk cloud auth' to link a LiveKit Cloud project.","'lk agent init my-agent --template agent-starter-python' creates the project and .env.local with credentials.","'uv sync', then 'lk agent dev' to run locally and talk to it from the Agents Playground or console mode.","'lk agent create' to deploy the agent to LiveKit Cloud; attach a phone number or SIP trunk with a dispatch rule for phone calls."],"endpoint":"wss://<project>.livekit.cloud","auth":"LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET (JWT access tokens for clients)","snippet_lang":"python","snippet":"# agent.py - shape per LiveKit voice AI quickstart (Oct 2026); model ids from the docs example\nfrom dotenv import load_dotenv\nfrom livekit import agents\nfrom livekit.agents import AgentServer, AgentSession, Agent, inference, TurnHandlingOptions\n\nload_dotenv(\".env.local\")\n\nclass Assistant(Agent):\n    def __init__(self):\n        super().__init__(instructions=\"You are a concise, friendly voice assistant.\")\n\nserver = AgentServer()\n\n@server.rtc_session(agent_name=\"my-agent\")\nasync def my_agent(ctx: agents.JobContext):\n    session = AgentSession(\n        stt=inference.STT(model=\"assemblyai/universal-3-6-pro\", language=\"en\"),\n        llm=inference.LLM(model=\"google/gemma-4-31b-it\"),\n        tts=inference.TTS(model=\"fishaudio/s2.1-pro\", voice=\"<voice-id>\"),\n        turn_handling=TurnHandlingOptions(turn_detection=inference.TurnDetector()),\n    )\n    await session.start(room=ctx.room, agent=Assistant())\n    await session.generate_reply(instructions=\"Greet the user and offer your assistance.\")\n\nif __name__ == \"__main__\":\n    agents.cli.run_app(server)  # python agent.py dev"},"warnings":[{"severity":"medium","title":"Agent session minutes are only the hosting fee","detail":"$0.01/min covers running your agent; STT, LLM and TTS are billed separately (via LiveKit Inference credits or your own provider keys), and phone minutes are billed again."},{"severity":"medium","title":"Free inference credit is tiny","detail":"Build includes $2.50 of inference (~50 minutes). After that, calls fail or need your own provider keys unless you upgrade."},{"severity":"medium","title":"Data region is immutable","detail":"Project data region (US or EU) is chosen at project creation and cannot be changed. Telephony, storage and third-party plugins each need their own regional configuration for GDPR setups."},{"severity":"low","title":"API churn between versions","detail":"The quickstart moved to AgentServer/@server.rtc_session and TurnHandlingOptions; older WorkerOptions/entrypoint examples on blogs may not match the current SDK. Pin versions."},{"severity":"low","title":"Turn detector model license","detail":"The Agents framework is Apache-2.0, but LiveKit's turn detection models are under the separate LiveKit Model License. Read it before self-hosting commercially."},{"severity":"low","title":"HIPAA status not confirmed","detail":"Official security page lists SOC 2 Type II and GDPR; a HIPAA/BAA statement was not confirmed in this research. Ask sales before handling PHI."}],"best_for":"Developers who want full code control of the voice pipeline with a managed WebRTC/telephony layer, or a path to self-hosting.","open_source":true,"self_hostable":true,"compliance":"SOC 2 Type II (Security, Availability, Confidentiality) and GDPR per LiveKit security page; HIPAA/BAA not verified.","docs":[{"label":"Pricing","url":"https://livekit.com/pricing"},{"label":"Voice AI quickstart","url":"https://docs.livekit.io/agents/start/voice-ai-quickstart/"},{"label":"Agents GitHub","url":"https://github.com/livekit/agents"},{"label":"Data residency","url":"https://docs.livekit.io/deploy/admin/regions/data-residency"}],"sources":["https://livekit.com/pricing","https://docs.livekit.io/agents/start/voice-ai-quickstart/","https://github.com/livekit/agents","https://livekit.com/security/overview","https://docs.livekit.io/deploy/admin/regions/data-residency"],"confidence":"high","unverified":"Scale plan concurrency (table on pricing page is ambiguous); HIPAA/BAA availability; exact model ids change frequently.","cat":"platforms","kind":"framework","verified_at":"2026-10-10","short":"LiveKit","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":20,"free_tier":true,"free_credit_usd":2.5,"entry_plan_usd_month":50,"hipaa":null,"soc2":true,"gdpr_eu":true,"webrtc":true,"websocket":null,"sip":true,"open_source":true,"self_hostable":true,"platform_fee_per_min":0.01,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":true,"no_code_builder":null,"recording":true,"noise_cancellation":true,"turn_detection_model":true,"_notes":"Free Build plan: 1,000 agent minutes, 5 concurrent, $2.50 inference credit. $0.01/min applies after included minutes. Ship concurrency 20; Scale listed up to 600 (ambiguous). HIPAA/BAA not confirmed.","_added":[]}},{"id":"pipecat-cloud","name":"Pipecat + Pipecat Cloud","vendor":"Daily","category":"framework","summary":"Pipecat is an open-source (BSD-2-Clause) Python framework for voice and multimodal agents with roughly 85 STT/LLM/TTS integrations; Pipecat Cloud is Daily's hosting service for Pipecat agents with Daily WebRTC and telephony built in.","status":"GA","models":[{"name":"Pipecat (open source)","status":"GA","notes":"Self-run anywhere; transports include Daily, LiveKit, SmallWebRTC, FastAPI/WebSocket, Vonage, WhatsApp."},{"name":"Pipecat Cloud agent-1x","status":"GA","notes":"0.5 vCPU, 1 GB: $0.01/min active, $0.0005/min reserved."},{"name":"Pipecat Cloud agent-2x","status":"GA","notes":"1 vCPU, 2 GB: $0.02/min active, $0.001/min reserved."},{"name":"Pipecat Cloud agent-3x","status":"GA","notes":"1.5 vCPU, 3 GB: $0.03/min active, $0.0015/min reserved."}],"transports":["WebRTC","WebSocket","SIP","PSTN"],"audio":{"input":"Daily WebRTC (Opus), WebSocket serializers for Twilio/Telnyx/Plivo/Vonage (8 kHz mu-law or PCM)","output":"Same as input"},"languages":"Depends on chosen providers.","latency":"No vendor number verified.","features":["turn detection (Smart Turn open model, Silero VAD)","noise cancellation (Krisp VIVA)","telephony (Daily SIP/PSTN, Twilio and others via WebSocket)","multi-provider plug-ins (~85 services)","recording","metrics","reserved warm instances","speech-to-speech services"],"pricing":{"model":"usage","items":[{"what":"Agent hosting agent-1x","price":"$0.01","unit":"per active minute","notes":"Compute only; LLM/STT/TTS billed by your providers (bring your own keys)."},{"what":"Reserved (always-on) agent-1x","price":"$0.0005","unit":"per reserved minute","notes":"Billed every minute until you turn it off; avoids cold starts."},{"what":"Daily WebRTC voice (1:1 on Pipecat Cloud)","price":"$0.001 listed, marked Free","unit":"per minute","notes":"Page labels 1:1 voice sessions free; voice+video $0.004/participant minute."},{"what":"Daily PSTN dial-in/out (SIP included)","price":"$0.018","unit":"per minute","notes":"Transcoding adds charges."},{"what":"Daily SIP dial-in/out","price":"$0.003-$0.02","unit":"per minute","notes":""},{"what":"SIP REFER transfer","price":"$0.20","unit":"per event","notes":"Minutes keep accruing until the handoff completes."},{"what":"Krisp VIVA noise cancellation","price":"Free to 10k min, then $0.0015","unit":"per minute","notes":"Per month."},{"what":"Recording (audio only)","price":"$0.005","unit":"per minute","notes":"Storage $0.003/min."}],"est_per_minute_usd":{"low":0.03,"high":0.1,"basis":"$0.01 agent-1x hosting + ~$0.005-0.01 STT + ~$0.002-0.02 LLM + ~$0.015-0.05 TTS (provider list prices, billed by providers) + $0 for 1:1 WebRTC or $0.018 for Daily PSTN. Not a vendor figure; it is an assembly of the listed rates with typical provider prices."},"free_tier":"Page says 'Start building for free'; 1:1 Daily WebRTC voice labelled free; Krisp free up to 10k min/month. No explicit credit amount found.","source":"https://www.daily.co/pricing/pipecat-cloud/"},"limits":["Pricing page says 'unlimited concurrency'","Cold starts unless you keep min_agents / reserved instances warm"],"regions":"Not verified in this research.","setup":{"steps":["Python 3.11+ and uv. Get Deepgram, OpenAI and Cartesia keys (quickstart stack).","Run the quickstart bot locally with 'uv run bot.py' and open the local WebRTC client.","Install the CLI: uv tool install \"pipecat-ai[cli]\"; then 'pipecat cloud auth login'.","'pipecat cloud secrets set <secret-set> --file .env', configure pcc-deploy.toml (agent_name, secret_set, scaling.min_agents), then 'pipecat cloud deploy' (cloud build, no registry needed).","Start sessions via the REST API or Python SDK; attach Daily phone numbers or SIP for calls."],"endpoint":"Pipecat Cloud REST API / Daily rooms","auth":"Pipecat Cloud API key (CLI login); provider keys stored as secret sets","snippet_lang":"python","snippet":"# Core of the Pipecat quickstart bot.py (Oct 2026). Imports omitted: copy them from\n# docs.pipecat.ai quickstart, class names changed in recent releases\n# (PipelineTask/PipelineRunner -> PipelineWorker/WorkerRunner).\nstt = DeepgramSTTService(api_key=os.getenv(\"DEEPGRAM_API_KEY\"))\ntts = CartesiaTTSService(api_key=os.getenv(\"CARTESIA_API_KEY\"),\n    settings=CartesiaTTSService.Settings(voice=os.getenv(\"CARTESIA_VOICE_ID\")))\nllm = OpenAIResponsesLLMService(api_key=os.getenv(\"OPENAI_API_KEY\"),\n    settings=OpenAIResponsesLLMService.Settings(model=\"gpt-4.1\",\n        system_instruction=\"You are a brief, friendly voice assistant.\"))\n\ncontext = LLMContext()\nuser_agg, assistant_agg = LLMContextAggregatorPair(\n    context, user_params=LLMUserAggregatorParams(vad_analyzer=SileroVADAnalyzer()))\n\npipeline = Pipeline([transport.input(), stt, user_agg, llm, tts,\n                     transport.output(), assistant_agg])\nworker = PipelineWorker(pipeline, params=PipelineParams(enable_metrics=True))\nrunner = WorkerRunner(handle_sigint=runner_args.handle_sigint)\nawait runner.add_workers(worker)\nawait runner.run()\n\n# Deploy: pipecat cloud auth login\n#         pipecat cloud secrets set my-secrets --file .env\n#         pipecat cloud deploy"},"warnings":[{"severity":"medium","title":"Hosting price excludes all AI models","detail":"Pipecat Cloud bills compute and Daily transport only; you pay Deepgram/OpenAI/Cartesia (or others) directly. Bundled inference billing exists only for enterprise via sales."},{"severity":"medium","title":"Reserved instances bill around the clock","detail":"Warm 'reserved' agents avoid cold starts but bill every minute until switched off ($0.0005/min for agent-1x is about $21.60 per 30-day month per instance)."},{"severity":"medium","title":"SIP transfers keep the meter running","detail":"With SIP REFER the call stays anchored to Daily until the handoff completes, then a $0.20 event fee applies."},{"severity":"low","title":"Fast-moving API","detail":"Class names and the CLI changed in 2025-2026 (pcc -> 'pipecat cloud', PipelineTask -> PipelineWorker). Pin the pipecat-ai version and follow the docs for your version."},{"severity":"low","title":"Telephony audio via WebSocket is 8 kHz","detail":"When using Twilio/Telnyx/Plivo serializers, audio is usually 8 kHz mu-law; configure sample rates correctly in the transport params."}],"best_for":"Python teams that want maximum provider choice and open-source portability, with optional managed hosting.","open_source":true,"self_hostable":true,"compliance":"Not verified in this research for Pipecat Cloud (ask Daily for SOC 2/HIPAA terms).","docs":[{"label":"Pipecat Cloud pricing","url":"https://www.daily.co/pricing/pipecat-cloud/"},{"label":"Quickstart","url":"https://docs.pipecat.ai/pipecat/get-started/quickstart"},{"label":"Pipecat Cloud deploy guide","url":"https://docs.pipecat.ai/guides/deployment/pipecat-cloud"},{"label":"GitHub","url":"https://github.com/pipecat-ai/pipecat"}],"sources":["https://www.daily.co/pricing/pipecat-cloud/","https://docs.pipecat.ai/pipecat/get-started/quickstart","https://github.com/pipecat-ai/pipecat","https://docs.pipecat.ai/pipecat-cloud/guides/cloud-builds"],"confidence":"high","unverified":"Regions, compliance attestations, free credit amount; snippet imports deliberately omitted because module paths changed in recent releases.","cat":"platforms","kind":"framework","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":true,"open_source":true,"self_hostable":true,"platform_fee_per_min":0.01,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":true,"no_code_builder":null,"recording":true,"noise_cancellation":true,"turn_detection_model":true,"_notes":"Fee is agent-1x hosting ($0.02 / $0.03 for larger sizes); open-source Pipecat is free to self-host. Pricing page says unlimited concurrency. No explicit free credit amount.","_added":[]}},{"id":"cloudflare-realtime","name":"Cloudflare Realtime (SFU, TURN, RealtimeKit, Realtime Agents / voice)","vendor":"Cloudflare","category":"transport","summary":"Cloudflare's global WebRTC SFU and TURN service, RealtimeKit SDKs/UI for meetings, and a beta voice-agent stack that runs STT, LLM and TTS on Workers AI from inside Workers/Durable Objects.","status":"Beta","models":[{"name":"Realtime SFU + TURN","status":"GA","notes":"$0.05/GB egress after 1,000 GB free per month; ingress free."},{"name":"RealtimeKit","status":"GA (billing stated on docs)","notes":"$0.0005/min audio-only participant, $0.002/min audio/video participant. Older docs said free during beta; the current page states usage is charged."},{"name":"Voice agents (@cloudflare/voice on Agents SDK)","status":"Beta","notes":"withVoice(Agent) mixin, Workers AI Flux STT + Workers AI TTS, onTurn handler, React hook useVoiceAgent."},{"name":"Realtime Agents runtime","status":"Beta","notes":"Launch blog said free to use during the open beta; current status not re-verified."}],"transports":["WebRTC","WebSocket"],"audio":{"input":"WebRTC (Opus) or WebSocket PCM via the WebSocket adapter","output":"Same"},"languages":"Depends on Workers AI or third-party models (Deepgram, ElevenLabs via providers).","latency":"Vendor emphasises 330+ city edge network; no verified end-to-end number.","features":["global SFU/TURN","WebSocket adapter for AI audio","turn detection (Workers AI smart-turn, Flux end-of-speech)","multi-provider plug-ins via AI Gateway","recording/export (RealtimeKit)","conversation persistence in Durable Objects"],"pricing":{"model":"usage","items":[{"what":"SFU and TURN egress","price":"$0.05","unit":"per GB","notes":"First 1,000 GB/month free, shared. A voice stream is a few hundred KB per minute, so transport is usually negligible."},{"what":"RealtimeKit audio-only participant","price":"$0.0005","unit":"per minute","notes":"Type set by preset, not by what is actually sent."},{"what":"RealtimeKit audio/video participant","price":"$0.002","unit":"per minute","notes":""},{"what":"RealtimeKit export (audio only)","price":"$0.003","unit":"per minute","notes":"Recording/RTMP/HLS audio-only."},{"what":"Workers AI Deepgram Nova-3 streaming STT","price":"$0.0092","unit":"per audio minute","notes":"Workers AI list price (preview pricing page); billed in Neurons at $0.011 per 1,000."},{"what":"Workers AI Deepgram Aura-1 TTS","price":"$0.015","unit":"per 1k characters","notes":"Aura-2-en listed at $0.030 per 1k characters."},{"what":"Workers AI free allowance","price":"10,000 Neurons","unit":"per day","notes":""}],"est_per_minute_usd":{"low":0.015,"high":0.05,"basis":"Workers AI STT ~$0.009/min + TTS ~$0.01-0.03/min (assuming roughly 600-1,000 spoken characters per agent minute) + small LLM token cost + near-zero transport within the free 1 TB. Assumes the Realtime Agents runtime itself stays free/cheap; not a vendor figure. No PSTN included (Cloudflare does not sell phone minutes)."},"free_tier":"1,000 GB/month SFU+TURN egress; 10,000 Workers AI Neurons/day.","source":"https://developers.cloudflare.com/realtime/sfu/pricing/"},"limits":["No native PSTN numbers; bring a SIP/telephony provider","Voice agent package is Beta"],"regions":"Global anycast edge (330+ cities per Cloudflare).","setup":{"steps":["Start from the cloudflare/agents-starter template and 'npm i @cloudflare/voice'.","In wrangler config add an 'ai' binding named AI, a Durable Object binding for the agent class with a new_sqlite_classes migration, and the nodejs_compat flag.","Implement onTurn(transcript, context) returning a text stream; use useVoiceAgent({ agent: 'MyVoiceAgent' }) in React for mic, status and startCall/endCall.","'npx wrangler deploy'."],"endpoint":"Your Worker URL (WebSocket to the Durable Object agent)","auth":"Cloudflare account / Workers AI binding; RealtimeKit uses app ID + API token","snippet_lang":"javascript","snippet":"// src/server.ts - shape per Cloudflare \"Voice agent\" docs (Beta). Check docs for exact STT/TTS wiring.\nimport { Agent } from \"agents\";\nimport { withVoice, WorkersAIFluxSTT, WorkersAITTS } from \"@cloudflare/voice\";\nimport { streamText } from \"ai\";\nimport { createWorkersAI } from \"workers-ai-provider\";\n\nconst VoiceAgent = withVoice(Agent);\n\nexport class MyVoiceAgent extends VoiceAgent {\n  stt = new WorkersAIFluxSTT(this.env.AI);   // end-of-speech aware STT\n  tts = new WorkersAITTS(this.env.AI);       // sentence-by-sentence TTS\n\n  async onCallStart(connection) {\n    await this.speak(\"Hi, how can I help?\");\n  }\n\n  async onTurn(transcript, context) {\n    const workersai = createWorkersAI({ binding: this.env.AI });\n    const result = streamText({\n      model: workersai(\"@cf/moonshotai/kimi-k2.6\"),\n      system: \"You are a voice assistant. Keep answers to one or two sentences.\",\n      messages: [...context.messages, { role: \"user\", content: transcript }],\n      abortSignal: context.signal, // aborts on barge-in\n    });\n    return result.textStream;\n  }\n}"},"warnings":[{"severity":"medium","title":"Voice agent stack is Beta","detail":"@cloudflare/voice and the Realtime Agents runtime are Beta; APIs and pricing may change. The docs example has already been revised (different compatibility dates)."},{"severity":"medium","title":"RealtimeKit billing status changed","detail":"Older docs said RealtimeKit was free during beta; the current pricing page states usage is charged. Check the calculator before assuming free."},{"severity":"medium","title":"No phone numbers","detail":"Cloudflare provides WebRTC/WebSocket transport, not PSTN. For phone agents you need Twilio/Telnyx/Plivo or another SIP provider."},{"severity":"low","title":"Participant type is set by preset","detail":"RealtimeKit bills audio/video rate if the preset meeting type is video, even if users only send audio."},{"severity":"low","title":"Workers AI model menu is narrower","detail":"Third-party models via AI Gateway are billed by those providers; only Workers AI models land on the Cloudflare bill."}],"best_for":"Edge-native web voice apps already on Workers, and anyone needing cheap global WebRTC transport.","open_source":false,"self_hostable":false,"compliance":"Not verified for these products in this research.","docs":[{"label":"SFU pricing","url":"https://developers.cloudflare.com/realtime/sfu/pricing/"},{"label":"RealtimeKit pricing","url":"https://developers.cloudflare.com/realtime/realtimekit/pricing"},{"label":"Voice agent example","url":"https://developers.cloudflare.com/agents/examples/voice-agent/"},{"label":"Realtime voice AI blog","url":"https://blog.cloudflare.com/cloudflare-realtime-voice-ai"},{"label":"Workers AI pricing","url":"https://developers.cloudflare.com/workers-ai/platform/pricing/"}],"sources":["https://developers.cloudflare.com/realtime/sfu/pricing/","https://developers.cloudflare.com/realtime/realtimekit/pricing","https://developers.cloudflare.com/agents/examples/voice-agent/","https://blog.cloudflare.com/cloudflare-realtime-voice-ai","https://developers.cloudflare.com/workers-ai/platform/pricing/"],"confidence":"medium","unverified":"Whether the Realtime Agents runtime is still free; Workers AI per-model prices came from a preview build of the pricing page; Flux STT price on Workers AI not found; exact property names for wiring STT/TTS in the snippet.","cat":"platforms","kind":"transport","verified_at":"2026-10-10","short":"Cloudflare Realtime","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"platform_fee_per_min":null,"all_in":false,"phone_numbers":false,"byo_llm":true,"byo_keys":true,"no_code_builder":null,"recording":true,"noise_cancellation":null,"turn_detection_model":true,"_notes":"Voice agent stack is Beta. SFU/TURN billed per GB ($0.05 after 1 TB/month free), RealtimeKit $0.0005/min per audio participant; agents runtime said free in beta (not re-verified). No PSTN numbers.","_added":[]}},{"id":"agora-conversational-ai","name":"Agora Conversational AI Engine","vendor":"Agora","category":"platform","summary":"Agora's hosted voice-agent engine that joins an Agora RTC channel (or SIP/PSTN call) and runs ASR, LLM and TTS for you, configured by REST; built on the open-source TEN framework.","status":"GA","models":[{"name":"Conversational AI Engine Audio Task","status":"GA","notes":"$0.10/min list price; includes Agora-managed ASR/LLM/TTS from a select list (ASR: ARES, Deepgram nova-2/3; LLM: GPT-4o-mini, GPT-4.1-mini, GPT-5-nano, GPT-5-mini; TTS: MiniMax, OpenAI TTS-1). BYOK users pay the same rate."},{"name":"Agent Studio","status":"GA","notes":"No-code UI for deploying agents, no separate price listed."},{"name":"Voice AI agents over SIP/PSTN","status":"GA","notes":"Billed on answered duration, rounded up to the next full minute."}],"transports":["WebRTC","SIP","PSTN"],"audio":{"input":"Agora RTC (SD-RTN) audio; SIP/PSTN","output":"Same"},"languages":"Depends on chosen ASR/TTS vendors.","latency":"No verified vendor number.","features":["managed ASR/LLM/TTS bundle","BYO LLM endpoint (OpenAI-compatible URL)","interruption handling","telephony","avatars (separate)","Agent Studio no-code"],"pricing":{"model":"per-minute","items":[{"what":"Audio Task (agent minute)","price":"$0.10","unit":"per minute","notes":"Includes select managed ASR, LLM and TTS models. Same price with your own keys. Bring-your-own-model pricing via sales."},{"what":"Audio RTC for each user","price":"$0.00099","unit":"per minute","notes":"Shown separately in Agora's billing example."},{"what":"Free minutes","price":"300 minutes","unit":"one-time","notes":"First 300 minutes free."}],"est_per_minute_usd":{"low":0.101,"high":0.12,"basis":"$0.10 agent task + $0.00099 RTC per user, using included managed models. High end allows for SIP/PSTN carrier minutes, which were not priced in the sources found."},"free_tier":"First 300 minutes free.","source":"https://docs.agora.io/en/conversational-ai/overview/pricing"},"limits":["Concurrency limits not published in sources reviewed","Managed-model list is limited; others need BYOK or BYOM (sales)"],"regions":"Agora SD-RTN global network.","setup":{"steps":["Create an Agora project (App ID + certificate) and enable Conversational AI in the console.","Have your client app join an RTC channel with a token.","POST to the join endpoint with channel, agent token, UIDs and asr/llm/tts config; store the returned agent_id.","Call the leave endpoint with agent_id when done (idle_timeout also stops it)."],"endpoint":"POST https://api.agora.io/api/conversational-ai-agent/v2/projects/{appid}/join","auth":"HTTP Basic (base64 of Customer ID:Customer Secret)","snippet_lang":"javascript","snippet":"// Start an Agora Conversational AI agent in an existing RTC channel (Node 18+)\nconst appId = process.env.AGORA_APP_ID;\nconst basic = Buffer.from(process.env.AGORA_CUSTOMER_ID + \":\" + process.env.AGORA_CUSTOMER_SECRET).toString(\"base64\");\n\nconst res = await fetch(\"https://api.agora.io/api/conversational-ai-agent/v2/projects/\" + appId + \"/join\", {\n  method: \"POST\",\n  headers: { Authorization: \"Basic \" + basic, \"Content-Type\": \"application/json\" },\n  body: JSON.stringify({\n    name: \"support-agent-\" + Date.now(),\n    properties: {\n      channel: \"room-123\",\n      token: process.env.AGENT_RTC_TOKEN,\n      agent_rtc_uid: \"0\",\n      remote_rtc_uids: [\"1002\"],\n      idle_timeout: 120,\n      // asr / llm / tts blocks: use credential_mode \"managed\" for Agora-billed models,\n      // or vendor + your own key. See the join API reference for field names per vendor.\n      llm: { system_messages: [{ role: \"system\", content: \"You are a helpful voice assistant.\" }],\n             greeting_message: \"Hi! How can I help?\" }\n    }\n  })\n});\nconsole.log(await res.json()); // { agent_id, create_ts, status: \"RUNNING\" }"},"warnings":[{"severity":"high","title":"Old pricing still circulates","detail":"Third-party pages and older Agora docs quote about $0.0265/min (Audio Basic $0.0099 + ARES ASR $0.0166) with LLM/TTS billed separately. The current pricing doc lists $0.10/min including select managed models. Check which model you are on before budgeting."},{"severity":"medium","title":"BYOK does not lower the price","detail":"Bringing your own ASR/LLM/TTS keys still costs $0.10/min, so you would pay twice (Agora + your provider)."},{"severity":"medium","title":"Phone calls round up to the full minute","detail":"SIP/PSTN agent calls bill answered duration rounded up to the next minute; ring time and missed calls are free."},{"severity":"low","title":"Token and UID plumbing","detail":"The agent needs its own RTC token and UID, and remote_rtc_uids must match the user, or the agent hears nobody."},{"severity":"low","title":"Cloud Proxy minimum is separate","detail":"Agora Cloud Proxy (for firewalled enterprise networks) has a separate monthly minimum starting at $500."}],"best_for":"Apps already on Agora RTC (including Asia-Pacific audiences) that want a single per-minute price including models.","open_source":false,"self_hostable":false,"compliance":"Not verified in this research.","docs":[{"label":"Pricing doc","url":"https://docs.agora.io/en/conversational-ai/overview/pricing"},{"label":"Pricing page","url":"https://www.agora.io/en/pricing/agora-conversational-ai-platform"},{"label":"Join API","url":"https://docs.agora.io/en/conversational-ai/rest-api/agent/join"}],"sources":["https://docs.agora.io/en/conversational-ai/overview/pricing","https://www.agora.io/en/pricing/agora-conversational-ai-platform","https://docs.agora.io/en/conversational-ai/rest-api/agent/join"],"confidence":"medium","unverified":"SIP/PSTN carrier rates, concurrency limits, compliance, exact asr/tts field schema per vendor.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","short":"Agora ConvoAI","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.1,"all_in":true,"phone_numbers":true,"byo_llm":true,"byo_keys":true,"no_code_builder":true,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"$0.10/min includes select managed ASR/LLM/TTS; same price with your own keys. First 300 minutes free. Plus $0.00099/min RTC per user.","_added":[]}},{"id":"ten-framework","name":"TEN Framework","vendor":"Agora / TEN community","category":"framework","summary":"Open-source framework for realtime multimodal conversational agents backed by Agora, with a visual graph designer (TMAN Designer), agent examples, and the TEN VAD and TEN Turn Detection models.","status":"GA","models":[{"name":"TEN Framework","status":"GA","notes":"Apache 2.0 with additional restrictions (packages directory is plain Apache 2.0)."},{"name":"TEN VAD","status":"GA","notes":"Lightweight low-latency streaming voice activity detector."},{"name":"TEN Turn Detection","status":"GA","notes":"Semantic end-of-turn model for full-duplex dialogue, English and Chinese; Apache 2.0 with additional restrictions."}],"transports":["WebRTC","SIP"],"audio":{"input":"Agora RTC by default; other extensions possible","output":"Same"},"languages":"Depends on providers; turn detection supports English and Chinese.","latency":"No verified number.","features":["visual graph designer","turn detection","VAD","multi-provider plug-ins (extensions)","avatars","SIP call example"],"pricing":{"model":"free/open","items":[{"what":"Framework","price":"$0","unit":"license","notes":"You pay Agora RTC (or other transport) and every model provider separately."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Depends entirely on chosen transport and providers; no vendor figure."},"free_tier":"Open source.","source":"https://github.com/TEN-framework/ten-framework"},"limits":["You operate everything: scaling, monitoring, TURN"],"regions":"Self-hosted.","setup":{"steps":["Clone github.com/TEN-framework/ten-framework and follow the ai_agents README (Docker, Node.js 18 prerequisite).","Fill in .env with Agora App ID and provider keys.","Start the dev stack, open TMAN Designer at localhost:49483, pick an example graph (voice assistant, SIP call) and edit extension properties."],"endpoint":"Self-hosted","auth":"Provider keys in .env","snippet_lang":"python","snippet":"# TEN is configured as a graph (JSON) of extensions rather than a short script.\n# Start from an example in the repo:\n#   git clone https://github.com/TEN-framework/ten-framework\n#   cd ten-framework/ai_agents   # follow README: copy .env.example to .env, add keys\n#   # start the dev containers as the README describes, then open\n#   # http://localhost:49483  (TMAN Designer) and select the voice-assistant graph"},"warnings":[{"severity":"medium","title":"License has extra restrictions","detail":"Both the framework and the turn-detection model are 'Apache 2.0 with additional restrictions'. Read LICENSE before commercial use."},{"severity":"medium","title":"Smaller ecosystem","detail":"Fewer integrations, docs and community answers than Pipecat or LiveKit Agents."},{"severity":"low","title":"No managed hosting from the project","detail":"Hosted use is via Agora's Conversational AI Engine (priced separately)."},{"severity":"low","title":"Releases not clearly versioned","detail":"The GitHub releases section was empty when checked; pin a commit."}],"best_for":"Teams on Agora who want to self-host and customise the same stack Agora's engine is built on, or who need the TEN VAD/turn models.","open_source":true,"self_hostable":true,"compliance":"N/A (self-hosted).","docs":[{"label":"GitHub","url":"https://github.com/TEN-framework/ten-framework"},{"label":"Agora TEN page","url":"https://www.agora.io/en/open-source/ten-framework"},{"label":"TEN Turn Detection","url":"https://github.com/TEN-framework/ten-turn-detection"}],"sources":["https://github.com/TEN-framework/ten-framework","https://www.agora.io/en/open-source/ten-framework","https://github.com/TEN-framework/ten-turn-detection","https://www.agora.io/en/blog/making-voice-ai-agents-more-human-with-ten-vad-and-turn-detection"],"confidence":"medium","unverified":"Exact startup commands (README changes often); latest version.","cat":"platforms","kind":"framework","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":true,"open_source":true,"self_hostable":true,"platform_fee_per_min":0,"all_in":false,"phone_numbers":false,"byo_llm":true,"byo_keys":true,"no_code_builder":true,"recording":null,"noise_cancellation":null,"turn_detection_model":true,"_notes":"Open source (Apache 2.0 with additional restrictions). Turn detection model covers English and Chinese only. Visual graph designer (TMAN). SIP via example extension, no numbers.","_added":[]}},{"id":"stream-video-vision-agents","name":"Stream Video & Audio + Vision Agents","vendor":"Stream (GetStream)","category":"transport","summary":"Stream's managed WebRTC edge network for video/audio calls, plus Vision Agents, an open-source (Apache-2.0) Python framework for voice and video AI agents that uses Stream as its default transport.","status":"GA","models":[{"name":"Build (free)","status":"GA","notes":"$100 free credits; page equates this to 333,000 audio participant minutes for calls."},{"name":"Pay-as-you-go","status":"GA","notes":"Audio $0.30 per 1,000 participant minutes for calls."},{"name":"Vision Agents (open source)","status":"GA","notes":"Python, Apache-2.0; plug-ins for OpenAI, Gemini Realtime, ElevenLabs, Deepgram and others; edge interface can target other WebRTC infra."}],"transports":["WebRTC"],"audio":{"input":"WebRTC","output":"WebRTC"},"languages":"Depends on providers.","latency":"No verified number.","features":["noise cancellation (add-on)","recording","transcription/captions (add-on)","multi-provider plug-ins (Vision Agents)","video processors (e.g. YOLO pose)"],"pricing":{"model":"usage","items":[{"what":"Audio participant minutes (calls)","price":"$0.30","unit":"per 1,000 participant minutes","notes":"An agent plus one user is 2 participants, about $0.0006/min."},{"what":"Noise cancellation","price":"$0.30","unit":"per 1,000 participant minutes","notes":""},{"what":"Transcription / closed captions","price":"$8.00","unit":"per 1,000 call minutes","notes":""},{"what":"Audio recording","price":"$1.50","unit":"per 1,000 call minutes","notes":""}],"est_per_minute_usd":{"low":0.02,"high":0.08,"basis":"Transport is about $0.0006/min for agent + user; the rest is your STT/LLM/TTS or realtime model bill from providers. Not a vendor figure."},"free_tier":"$100 free credits (Build).","source":"https://getstream.io/video/pricing/"},"limits":["No PSTN in sources reviewed","HIPAA via sales"],"regions":"Global edge (not detailed).","setup":{"steps":["Create a Stream app (API key + secret).","uv add \"vision-agents[getstream, openai, elevenlabs, deepgram]\".","Define an Agent with edge=getstream.Edge(), an agent_user, instructions and an llm (or STT/LLM/TTS), then have it join a Stream call your client app opens."],"endpoint":"Stream Video API","auth":"STREAM_API_KEY / STREAM_API_SECRET; provider keys","snippet_lang":"python","snippet":"# Vision Agents (shape from the GitHub README example; see README for the full runnable file)\n# uv add \"vision-agents[getstream, openai]\"\nfrom vision_agents.core import Agent, User          # verify import path in README\nfrom vision_agents.plugins import getstream, openai\n\nagent = Agent(\n    edge=getstream.Edge(),\n    agent_user=User(name=\"Assistant\"),\n    instructions=\"You are a friendly voice assistant. Keep replies short.\",\n    llm=openai.Realtime(),   # or separate stt=..., llm=..., tts=...\n)\n# Then create/join a Stream call from your app and have the agent join it."},"warnings":[{"severity":"medium","title":"Transport is cheap, models are not","detail":"Stream's audio rate is a fraction of a cent per minute; nearly all cost is the AI providers you plug in."},{"severity":"medium","title":"Snippet imports not verified","detail":"README only shows a partial example; check the docs for exact import paths and the call-join helper."},{"severity":"low","title":"No telephony","detail":"Stream is a WebRTC network; phone agents need a separate SIP/PSTN provider."},{"severity":"low","title":"Captions billed per call minute","detail":"Transcription add-on is $8 per 1,000 call minutes, separate from any STT you run for the agent."}],"best_for":"In-app voice/video agents, especially with vision (camera) input.","open_source":true,"self_hostable":true,"compliance":"HIPAA: contact sales (per pricing page). Others not verified.","docs":[{"label":"Video pricing","url":"https://getstream.io/video/pricing/"},{"label":"Vision Agents","url":"https://getstream.io/vision-agents/"},{"label":"GitHub","url":"https://github.com/GetStream/Vision-Agents"}],"sources":["https://getstream.io/video/pricing/","https://getstream.io/vision-agents/","https://github.com/GetStream/Vision-Agents"],"confidence":"medium","unverified":"Vision Agents import paths and join API; regions; SOC 2.","cat":"platforms","kind":"transport","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":100,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":false,"open_source":true,"self_hostable":true,"platform_fee_per_min":0.0006,"all_in":false,"phone_numbers":false,"byo_llm":true,"byo_keys":true,"no_code_builder":null,"recording":true,"noise_cancellation":true,"turn_detection_model":null,"_notes":"Fee is $0.30 per 1,000 participant minutes for agent plus one user. Noise cancellation and recording are paid add-ons. HIPAA via sales. No telephony.","_added":[]}},{"id":"vonage-ai-connectors","name":"Vonage Voice/Video API AI connectors (Audio Connector) + AI Studio","vendor":"Vonage (Ericsson)","category":"telephony","summary":"Vonage's Voice and Video APIs can stream call audio over WebSocket (Audio Connector, 8/16/24 kHz PCM) or WebRTC (Video Connector) to your own AI pipeline, with a Pipecat serializer; AI Studio is the low-code agent builder priced per session.","status":"GA","models":[{"name":"Audio Connector Server SDK","status":"GA","notes":"Python; WebSocket PCM 8/16/24 kHz; bidirectional flag lets you inject audio back; works with Voice API and Video API."},{"name":"Video Connector Server SDK","status":"GA","notes":"WebRTC, audio+video, Video API only."},{"name":"AI Studio","status":"GA","notes":"Session-volume tiers (10K-300K sessions/month) plus per-request AI charges; voice minutes billed at standard Voice API rates."}],"transports":["WebSocket","WebRTC","SIP","PSTN"],"audio":{"input":"Linear PCM 8/16/24 kHz over WebSocket","output":"Same (bidirectional)"},"languages":"Depends on your models.","latency":"No verified number.","features":["telephony","WebSocket audio streaming","Pipecat serializer","per-second billing on Voice API (vendor claim)","low-code builder (AI Studio)"],"pricing":{"model":"usage","items":[{"what":"Voice API minutes","price":"Not verified","unit":"per minute (billed per second)","notes":"vonage.com pricing page returned 403 to the research tool; rates vary by country."},{"what":"AI Studio NLU / Knowledge AI","price":"about $0.007-$0.008","unit":"per request","notes":"Sources disagree ($0.00803, $0.007, $0.0073); Knowledge AI $0.0073/request."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Could not verify current Vonage per-minute voice rates. Vapi's pricing page lists Vonage telephony at $0.00814/min as a pass-through reference."},"free_tier":"Not verified.","source":"https://developer.vonage.com/en/video/ai-connectors/overview"},"limits":["AI Studio fair use: average max 50 events per session"],"regions":"Global.","setup":{"steps":["Create a Vonage application with Voice (and/or Video) capability and a number.","For phone calls, answer with an NCCO 'connect' action to a websocket endpoint, or use the Audio Connector SDK for Video sessions.","Run your STT/LLM/TTS pipeline on the WebSocket (Pipecat has a Vonage serializer) and stream PCM back."],"endpoint":"Voice API answer_url returning NCCO; WebSocket endpoint you host","auth":"Application ID + private key (JWT)","snippet_lang":"javascript","snippet":"// Express answer_url: connect an inbound Vonage call to your AI WebSocket (16 kHz linear PCM)\nimport express from \"express\";\nconst app = express();\napp.get(\"/answer\", (req, res) => {\n  res.json([\n    { action: \"talk\", text: \"This call is with an AI assistant and may be recorded.\" },\n    { action: \"connect\",\n      endpoint: [{ type: \"websocket\", uri: \"wss://ai.example.com/vonage\",\n                   \"content-type\": \"audio/l16;rate=16000\",\n                   headers: { caller: req.query.from } }] }\n  ]);\n});\napp.listen(3000);"},"warnings":[{"severity":"medium","title":"Pricing not verified","detail":"Vonage's voice pricing page could not be fetched in this research; confirm rates in the dashboard."},{"severity":"medium","title":"AI Studio priced per session, not per minute","detail":"Sessions plus per-request NLU/Knowledge AI charges plus standard voice minutes; professional services packs start at $3,300/month."},{"severity":"low","title":"NCCO snippet from general knowledge","detail":"The websocket connect NCCO shape is long-standing but was not re-verified against current docs here."},{"severity":"low","title":"Bring your own AI stack","detail":"Connectors only move audio; all STT/LLM/TTS costs and latency are yours."}],"best_for":"Teams already on Vonage telephony or Video API who want to plug in their own AI pipeline.","open_source":false,"self_hostable":false,"compliance":"Not verified in this research.","docs":[{"label":"AI connectors overview","url":"https://developer.vonage.com/en/video/ai-connectors/overview"},{"label":"Audio Connector SDK","url":"https://developer.vonage.com/en/video/guides/audio-connector-sdk"},{"label":"AI Studio pricing","url":"https://www.vonage.com/communications-apis/ai-studio/pricing/"}],"sources":["https://developer.vonage.com/en/video/ai-connectors/overview","https://developer.vonage.com/en/blog/introducing-audio-connector-sdk-and-pipecat-serializer-for-ai-audio-apps","https://www.vonage.com/communications-apis/ai-studio/pricing/"],"confidence":"low","unverified":"All per-minute voice rates; AI Studio tier prices; NCCO websocket fields.","cat":"platforms","kind":"telephony","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":null,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":true,"no_code_builder":true,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Voice API per-minute rates could not be verified. AI Studio (low-code) is priced per session plus per-request NLU charges.","_added":[]}},{"id":"twilio-conversationrelay","name":"Twilio ConversationRelay","vendor":"Twilio","category":"telephony","summary":"Managed voice layer on Twilio calls: Twilio handles speech recognition (Deepgram or Google), text-to-speech (ElevenLabs, Google, Amazon), interruptions and turn-taking, and sends you text over a WebSocket so you only run the LLM logic.","status":"GA","models":[{"name":"ConversationRelay","status":"GA","notes":"$0.07/min on the US voice pricing page, plus normal voice minutes. Requires onboarding and accepting Twilio's AI addendum."}],"transports":["PSTN","SIP","WebSocket"],"audio":{"input":"Phone call; Twilio does STT and sends text to you","output":"You send text tokens; Twilio does TTS"},"languages":"Many via Google/Deepgram STT and ElevenLabs/Google/Amazon voices; 'multi' auto-detect requires Deepgram + ElevenLabs.","latency":"No verified number.","features":["turn detection (Deepgram Flux eotThreshold, speechTimeout)","interruptions (interruptible, interruptSensitivity)","telephony","DTMF","language switching","end session with handoffData","welcome greeting"],"pricing":{"model":"per-minute","items":[{"what":"ConversationRelay","price":"$0.07","unit":"per minute","notes":"Twilio says voice costs are calculated separately. Third-party sources describe the speech side (STT/TTS) as included; Twilio's page does not state it explicitly."},{"what":"Inbound local call","price":"$0.0085","unit":"per minute","notes":"Plus $1.15/month per local number."},{"what":"Outbound call (US/Canada)","price":"$0.014","unit":"per minute","notes":""},{"what":"Recording","price":"$0.0025","unit":"per minute","notes":"Storage $0.0005/min/month."},{"what":"Branded calling","price":"$0.12","unit":"per call","notes":"Optional."},{"what":"Answering machine detection","price":"$0.0075","unit":"per call","notes":"Optional."}],"est_per_minute_usd":{"low":0.08,"high":0.12,"basis":"$0.07 ConversationRelay + $0.0085 inbound or $0.014 outbound + your own LLM (roughly $0.002-0.03/min depending on model). Assumes STT/TTS are included in the $0.07 as third parties report."},"free_tier":"Twilio trial credit only.","source":"https://www.twilio.com/en-us/voice/pricing/us"},"limits":["url must be wss://","speechTimeout 600-5000 ms","Invalid provider/voice combos disconnect the call","Some STT/TTS providers are not PCI compliant; keep PCI data out of prompts and greetings"],"regions":"Twilio global; not region-pinned in sources reviewed.","setup":{"steps":["Complete ConversationRelay onboarding and accept the AI/ML addendum in the Twilio Console.","Buy a number and point its voice webhook to a TwiML endpoint that returns <Connect><ConversationRelay url=\"wss://...\">.","Run a WebSocket server: on 'setup' store call info, on 'prompt' send the caller text to your LLM and stream back 'text' tokens with last=true on the final one.","Handle 'interrupt' by cancelling the in-flight LLM stream."],"endpoint":"TwiML <Connect><ConversationRelay> + your wss:// server","auth":"Twilio Account SID / Auth Token (validate X-Twilio-Signature on webhooks)","snippet_lang":"javascript","snippet":"// Node: TwiML webhook + ConversationRelay WebSocket (message names per Twilio docs; verify schema)\nimport express from \"express\";\nimport { WebSocketServer } from \"ws\";\nconst app = express();\napp.post(\"/voice\", (req, res) => {\n  res.type(\"text/xml\").send(\n    '<Response><Connect><ConversationRelay url=\"wss://ai.example.com/relay\" ' +\n    'welcomeGreeting=\"Hi, this is an AI assistant. How can I help?\" /></Connect></Response>');\n});\nconst server = app.listen(3000);\nconst wss = new WebSocketServer({ server, path: \"/relay\" });\nwss.on(\"connection\", (ws) => {\n  ws.on(\"message\", async (raw) => {\n    const msg = JSON.parse(raw);\n    if (msg.type === \"setup\") console.log(\"call\", msg.callSid);\n    if (msg.type === \"prompt\") {\n      const reply = await askLLM(msg.voicePrompt);       // your LLM call\n      ws.send(JSON.stringify({ type: \"text\", token: reply, last: true }));\n    }\n    if (msg.type === \"interrupt\") { /* cancel in-flight LLM stream */ }\n  });\n});"},"warnings":[{"severity":"medium","title":"Two meters: $0.07/min plus the phone minutes","detail":"ConversationRelay is billed on top of inbound/outbound voice minutes and number rental; your LLM is billed by its provider."},{"severity":"medium","title":"Default STT changed in 2025","detail":"transcriptionProvider defaults to Deepgram, but accounts that used ConversationRelay before 12 Sep 2025 default to Google. Set it explicitly so behaviour does not differ between accounts."},{"severity":"medium","title":"Bad provider/voice combos hang up the call","detail":"An invalid transcriptionProvider/speechModel or ttsProvider/voice pair sends an error and disconnects. Test every voice you configure."},{"severity":"low","title":"Default for input during agent speech changed","detail":"reportInputDuringAgentSpeech default became 'none' in May 2025 (was 'any'); older tutorials assume otherwise."},{"severity":"low","title":"HIPAA and PCI","detail":"HIPAA use needs a signed BAA; keep card data out of welcomeGreeting, hints, parameters and handoffData because some providers are not PCI compliant."}],"best_for":"Teams that want to own the LLM/agent logic in plain text while Twilio runs speech and telephony.","open_source":false,"self_hostable":false,"compliance":"HIPAA eligible with a signed BAA (per TwiML docs); Twilio-wide SOC 2/GDPR programs not re-verified here.","docs":[{"label":"US voice pricing","url":"https://www.twilio.com/en-us/voice/pricing/us"},{"label":"ConversationRelay TwiML","url":"https://www.twilio.com/docs/voice/twiml/connect/conversationrelay"},{"label":"ConversationRelay overview","url":"https://www.twilio.com/docs/voice/conversationrelay"}],"sources":["https://www.twilio.com/en-us/voice/pricing/us","https://www.twilio.com/docs/voice/twiml/connect/conversationrelay","https://www.twilio.com/docs/voice/conversationrelay","https://twilio.com/en-us/products/conversational-ai/pricing"],"confidence":"medium","unverified":"Whether STT/TTS (especially ElevenLabs) are fully included in $0.07/min; exact WebSocket message field names (voicePrompt etc.).","cat":"platforms","kind":"telephony","verified_at":"2026-10-10","short":"Twilio ConversationRelay","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":true,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.07,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":null,"no_code_builder":null,"recording":true,"noise_cancellation":null,"turn_detection_model":true,"_notes":"$0.07/min reportedly includes STT and TTS (not stated explicitly by Twilio); you run the LLM. Voice minutes and numbers billed on top. Free tier is Twilio trial credit only. HIPAA needs a signed BAA.","_added":[]}},{"id":"twilio-media-streams","name":"Twilio Media Streams","vendor":"Twilio","category":"telephony","summary":"Raw call audio over WebSocket: bidirectional streams send you the caller's audio as base64 8 kHz mu-law and let you play audio back, so you can connect any STT/LLM/TTS or speech-to-speech model to a phone call.","status":"GA","models":[{"name":"Bidirectional stream (<Connect><Stream>)","status":"GA","notes":"One per call; inbound track only; blocks following TwiML until the socket closes."},{"name":"Unidirectional stream (<Start><Stream>)","status":"GA","notes":"Listen-only, up to 4 tracks per call shared with SIPREC, real-time transcription and AMD."}],"transports":["PSTN","SIP","WebSocket"],"audio":{"input":"audio/x-mulaw 8000 Hz mono, base64 in JSON 'media' events","output":"Same format sent back in 'media' events; 'clear' to flush on barge-in, 'mark' to track playback"},"languages":"Whatever your models support.","latency":"Not applicable (transport); your pipeline determines latency.","features":["telephony","raw audio access","barge-in via clear","playback marks","DTMF (bidirectional, inbound only)","works with OpenAI/Gemini realtime APIs"],"pricing":{"model":"per-minute","items":[{"what":"Media Streams","price":"$0.0044","unit":"per minute","notes":"Transport only."},{"what":"Inbound local / outbound US","price":"$0.0085 / $0.014","unit":"per minute","notes":"Plus $1.15/month per local number; toll-free inbound $0.022/min."},{"what":"SIP interface","price":"$0.004","unit":"per minute","notes":"If you bring calls in over SIP."}],"est_per_minute_usd":{"low":0.035,"high":0.12,"basis":"$0.0044 stream + $0.0085-$0.014 voice + your STT/LLM/TTS ($0.02-$0.10 depending on providers; speech-to-speech realtime models can exceed this)."},"free_tier":"Twilio trial credit only.","source":"https://www.twilio.com/en-us/voice/pricing/us"},"limits":["One bidirectional stream per call","Unidirectional: max 4 tracks per call (shared)","Stream resource API cannot start bidirectional streams"],"regions":"Twilio global.","setup":{"steps":["Point a Twilio number's voice webhook at TwiML returning <Connect><Stream url=\"wss://...\"/>.","Accept the WebSocket; read 'start' for streamSid, decode 'media' payloads (mu-law 8 kHz) into your STT or realtime model.","Send TTS audio back as mu-law 8 kHz base64 in 'media' events; send 'clear' when the caller interrupts."],"endpoint":"wss:// endpoint you host","auth":"Twilio credentials for webhooks; validate X-Twilio-Signature","snippet_lang":"javascript","snippet":"// Twilio bidirectional Media Stream skeleton\nimport express from \"express\";\nimport { WebSocketServer } from \"ws\";\nconst app = express();\napp.post(\"/voice\", (req, res) => res.type(\"text/xml\").send(\n  '<Response><Connect><Stream url=\"wss://ai.example.com/media\" /></Connect></Response>'));\nconst wss = new WebSocketServer({ server: app.listen(3000), path: \"/media\" });\nwss.on(\"connection\", (ws) => {\n  let streamSid;\n  ws.on(\"message\", (raw) => {\n    const m = JSON.parse(raw);\n    if (m.event === \"start\") streamSid = m.start.streamSid;\n    if (m.event === \"media\") {\n      const ulaw8k = Buffer.from(m.media.payload, \"base64\");\n      sendToSTTorRealtimeModel(ulaw8k);            // your pipeline\n    }\n  });\n  // when your TTS produces mu-law 8 kHz audio:\n  onAgentAudio((ulawChunk) => ws.send(JSON.stringify(\n    { event: \"media\", streamSid, media: { payload: ulawChunk.toString(\"base64\") } })));\n  onUserBargeIn(() => ws.send(JSON.stringify({ event: \"clear\", streamSid })));\n});"},"warnings":[{"severity":"medium","title":"8 kHz mu-law in and out","detail":"You must convert to/from mu-law 8 kHz. Sending PCM16 or 24 kHz audio unconverted produces noise. Many realtime model APIs accept g711_ulaw directly, which avoids resampling."},{"severity":"medium","title":"You own barge-in","detail":"Without sending 'clear' on interruption, Twilio keeps playing buffered audio after the caller starts talking."},{"severity":"medium","title":"<Connect><Stream> blocks the rest of the TwiML","detail":"Nothing after it runs until the socket closes; plan transfers via the REST API or by closing the stream and redirecting."},{"severity":"low","title":"Per-call track limits","detail":"Unidirectional streams, SIPREC, real-time transcription and AMD share a 4-track cap; exceeding it silently prevents the stream (stream-stopped callback)."}],"best_for":"Developers wiring their own pipeline or a speech-to-speech model (OpenAI Realtime, Gemini Live) to phone numbers.","open_source":false,"self_hostable":false,"compliance":"Twilio offers BAAs for HIPAA-eligible products; not re-verified for Media Streams here.","docs":[{"label":"Media Streams overview","url":"https://www.twilio.com/docs/voice/media-streams"},{"label":"US voice pricing","url":"https://www.twilio.com/en-us/voice/pricing/us"}],"sources":["https://www.twilio.com/docs/voice/media-streams","https://www.twilio.com/en-us/voice/pricing/us"],"confidence":"high","unverified":"Message field names in the snippet come from long-standing Twilio docs (WebSocket messages page not fetched in this session).","cat":"platforms","kind":"telephony","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.0044,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":true,"no_code_builder":null,"recording":true,"noise_cancellation":null,"turn_detection_model":false,"_notes":"Raw 8 kHz mu-law audio transport; you handle barge-in and turn-taking. Voice minutes billed on top. Free tier is Twilio trial credit only. BAA not re-verified for Media Streams.","_added":[]}},{"id":"telnyx-voice-ai","name":"Telnyx Voice AI (AI Assistants)","vendor":"Telnyx","category":"telephony","summary":"Carrier-owned voice-agent platform: Telnyx runs STT, TTS and orchestration on its own GPUs next to its telephony network for one engine rate, with LLM tokens and carrier minutes added on top.","status":"GA","models":[{"name":"Voice engine","status":"GA","notes":"$0.05/min covers orchestration (turn-taking, interruptions, tools, knowledge base retrieval), STT (Deepgram or Telnyx STT, hosted by Telnyx) and TTS (Telnyx Ultra/Natural/NaturalHD, Qwen3TTS, Inworld, Rime, Resemble, Murf)."},{"name":"LLM","status":"GA","notes":"Per-token: open models on Telnyx GPUs (Kimi 'typically ~$0.006/min') or managed OpenAI/Anthropic at their rates."},{"name":"Plans","status":"GA","notes":"Pay as you go ($0/mo, 500 concurrent calls), Growth ($2,000/mo minimum, 15% off voice/messaging/numbers), Enterprise ($5,000/mo minimum)."}],"transports":["PSTN","SIP","WebSocket","WebRTC"],"audio":{"input":"Phone/SIP; WebSocket to assistant","output":"Same"},"languages":"Not verified per voice; depends on STT/TTS chosen.","latency":"Vendor claims sub-500 ms end-to-end on marketing pages (vendor claim, not verified).","features":["turn detection","interruptions","tools/function calls","knowledge base","telephony (own carrier network)","no-code portal builder","MCP server","WebSocket assistant access"],"pricing":{"model":"per-minute","items":[{"what":"Voice AI engine (STT+TTS+orchestration)","price":"$0.05","unit":"per minute","notes":"No separate platform fee for the engine."},{"what":"LLM","price":"~$0.006 (Kimi on Telnyx GPUs) or provider rates","unit":"per minute equivalent","notes":"Billed per token."},{"what":"Inbound SIP trunking","price":"from $0.0032","unit":"per minute","notes":"Outbound SIP from $0.005/min."},{"what":"Voice API platform fee","price":"$0.002","unit":"per minute","notes":"On top of SIP trunking."}],"est_per_minute_usd":{"low":0.06,"high":0.09,"basis":"Telnyx's own 'realistic all-in' is about $0.06/min (example: ~$13,700/month at 230,000 minutes) with an open model on Telnyx GPUs. High end assumes a frontier OpenAI/Anthropic model and outbound calls. Excludes 60-second per-call rounding."},"free_tier":"No free usage credits listed; $0 monthly fee on pay as you go.","source":"https://telnyx.com/pricing/voice-ai"},"limits":["Pay as you go: 500 concurrent calls, 100 API requests/s","Per-call rounding to 60-second increments"],"regions":"Owned network with 100+ PoPs; regional inference (e.g. Paris) per Telnyx pages.","setup":{"steps":["Create a Telnyx account and API key; buy a number.","Create an assistant in the portal (no-code) or via POST /v2/ai/assistants with name, instructions and model.","Assign the assistant to the number (portal) and call it; or connect over WebSocket for web clients."],"endpoint":"POST https://api.telnyx.com/v2/ai/assistants","auth":"Authorization: Bearer <TELNYX_API_KEY>","snippet_lang":"javascript","snippet":"// Create a Telnyx AI Assistant (then attach it to a number in the portal)\nconst res = await fetch(\"https://api.telnyx.com/v2/ai/assistants\", {\n  method: \"POST\",\n  headers: { Authorization: \"Bearer \" + process.env.TELNYX_API_KEY,\n             \"Content-Type\": \"application/json\" },\n  body: JSON.stringify({\n    name: \"Front desk\",\n    instructions: \"You are the front desk for Acme Dental. Disclose you are an AI. Book appointments; keep replies short.\",\n    // model: optional; Telnyx applies a default if omitted (see API reference for voice/transcription objects)\n  }),\n});\nconsole.log(await res.json());"},"warnings":[{"severity":"medium","title":"60-second rounding","detail":"Telnyx says its estimates exclude per-call rounding to 60-second increments, so a 10-second call can bill as a full minute."},{"severity":"medium","title":"LLM not in the $0.05","detail":"Engine rate covers STT/TTS/orchestration only; LLM tokens and carrier minutes are extra. Prompt size drives LLM cost when using OpenAI/Anthropic."},{"severity":"medium","title":"No free credits","detail":"Unlike most platforms there is no free usage allowance; testing costs money from the first call."},{"severity":"low","title":"BAA terms unclear","detail":"Site lists HIPAA, SOC 2 Type II, PCI, ISO and GDPR, but an explicit BAA covering Voice AI was not found; get it in writing."}],"best_for":"High-volume phone agents where carrier + AI on one bill and low telephony cost matter.","open_source":false,"self_hostable":false,"compliance":"Site footer lists ISO, PCI, HIPAA, GDPR, SOC 2 Type II; Voice AI BAA scope not verified.","docs":[{"label":"Voice AI pricing","url":"https://telnyx.com/pricing/voice-ai"},{"label":"AI Assistants docs","url":"https://developers.telnyx.com/docs/inference/ai-assistants"},{"label":"Create assistant API","url":"https://developers.telnyx.com/api-reference/assistants/create-an-assistant"}],"sources":["https://telnyx.com/pricing/voice-ai","https://developers.telnyx.com/api-reference/assistants/create-an-assistant","https://developers.telnyx.com/docs/inference/ai-assistants","https://telnyx.com/products/llm-library"],"confidence":"medium","unverified":"Exact voice/transcription JSON fields; LLM per-token tables; latency claim.","cat":"platforms","kind":"telephony","verified_at":"2026-10-10","short":"Telnyx Voice AI","facts":{"latency_ms":500,"languages":null,"max_session_min":null,"concurrency":500,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":true,"gdpr_eu":true,"webrtc":true,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.05,"all_in":false,"phone_numbers":true,"byo_llm":null,"byo_keys":null,"no_code_builder":true,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Latency is a vendor marketing claim (sub-500 ms). $0.05/min engine includes Telnyx-hosted STT and TTS; LLM tokens and carrier minutes extra. 500 concurrent on pay as you go. 60-second per-call rounding. HIPAA listed but Voice AI BAA scope unverified.","_added":[]}},{"id":"plivo-audio-streaming","name":"Plivo Voice + Audio Streaming","vendor":"Plivo","category":"telephony","summary":"Low-cost telephony with bidirectional WebSocket audio streaming included in the call price, supporting 8 kHz mu-law or 16 kHz linear PCM, for connecting your own AI pipeline.","status":"GA","models":[{"name":"Audio Streaming (<Stream>)","status":"GA","notes":"Bidirectional option, keepCallAlive, contentType audio/x-mulaw;rate=8000 or audio/x-l16;rate=8000|16000; playAudio, checkpoint and clearAudio commands."}],"transports":["PSTN","SIP","WebSocket"],"audio":{"input":"mu-law 8 kHz or L16 8/16 kHz over WebSocket","output":"playAudio events with base64 audio, 8000 or 16000 Hz"},"languages":"Your models.","latency":"Not applicable (transport).","features":["telephony","audio streaming included","noise cancellation included","16 kHz streams","free call recording (storage free 90 days)","AMD at $0"],"pricing":{"model":"per-minute","items":[{"what":"Inbound local","price":"$0.0055","unit":"per minute","notes":"Audio streaming and noise cancellation included."},{"what":"Outbound local","price":"$0.0115","unit":"per minute","notes":""},{"what":"Browser SDK / SIP","price":"$0.0033","unit":"per minute","notes":""},{"what":"Local number","price":"$0.50","unit":"per month","notes":"Toll-free $1.00/month."},{"what":"Call recording","price":"$0.00","unit":"per minute","notes":"Storage free 90 days then $0.0004/min/month."}],"est_per_minute_usd":{"low":0.03,"high":0.1,"basis":"$0.0055-$0.0115 telephony + your own STT/LLM/TTS ($0.02-$0.09). Not a vendor figure."},"free_tier":"Not verified.","source":"https://www.plivo.com/voice/pricing/us/"},"limits":["Not published in sources reviewed"],"regions":"Global carrier coverage (not detailed).","setup":{"steps":["Buy a Plivo number and set its answer URL to an endpoint returning Plivo XML with <Stream>.","Accept the WebSocket; decode incoming audio; run your pipeline.","Send playAudio events back; use clearAudio on barge-in and checkpoint to know when playback finished."],"endpoint":"Answer URL returning XML; your wss:// server","auth":"Plivo Auth ID / Auth Token","snippet_lang":"javascript","snippet":"// Plivo answer URL: bidirectional 16 kHz stream to your AI server\nimport express from \"express\";\nconst app = express();\napp.all(\"/answer\", (req, res) => {\n  res.type(\"application/xml\").send(\n    '<Response>' +\n      '<Stream bidirectional=\"true\" keepCallAlive=\"true\" ' +\n      'contentType=\"audio/x-l16;rate=16000\">wss://ai.example.com/plivo</Stream>' +\n    '</Response>');\n});\napp.listen(3000);\n// On the socket, send audio back as:\n// { \"event\": \"playAudio\", \"media\": { \"contentType\": \"audio/x-l16\", \"sampleRate\": 16000, \"payload\": \"<base64>\" } }"},"warnings":[{"severity":"medium","title":"Default content type is inconsistent in docs","detail":"One page says the default is mu-law 8 kHz, the XML reference says L16 8 kHz. Always set contentType explicitly and match sampleRate in playAudio."},{"severity":"medium","title":"keepCallAlive defaults to false","detail":"If the stream ends the call hangs up unless keepCallAlive=\"true\"."},{"severity":"low","title":"Bidirectional limits audioTrack","detail":"With bidirectional=true, audioTrack cannot be outbound or both."},{"severity":"low","title":"AI agent product pricing not on this page","detail":"Plivo's own AI agent offering was not priced on the voice pricing page; check the AI Agents page or sales."}],"best_for":"Cheapest DIY phone transport for your own AI pipeline, with optional 16 kHz audio.","open_source":false,"self_hostable":false,"compliance":"Not verified in this research.","docs":[{"label":"US voice pricing","url":"https://www.plivo.com/voice/pricing/us/"},{"label":"Stream element","url":"https://docs.plivo.com/docs/voice/xml/the-stream-element"},{"label":"Audio streaming guide","url":"https://plivo.com/docs/voice-agents/audio-streaming/concepts/audio-streaming-guide"}],"sources":["https://www.plivo.com/voice/pricing/us/","https://docs.plivo.com/docs/voice/xml/the-stream-element","https://www.plivo.com/docs/voice/concepts/audio-streaming/"],"confidence":"medium","unverified":"Free trial terms, concurrency limits, compliance.","cat":"platforms","kind":"telephony","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":true,"no_code_builder":null,"recording":true,"noise_cancellation":true,"turn_detection_model":null,"_notes":"Audio streaming, noise cancellation and recording are included in call price (inbound local $0.0055/min). Free trial terms not verified.","_added":[]}},{"id":"signalwire-ai","name":"SignalWire AI Agent (SWML)","vendor":"SignalWire","category":"platform","summary":"Telephony platform with a built-in AI agent runtime defined declaratively in SWML: one per-minute rate covers realtime STT, the LLM and standard TTS, plus SWAIG function calling to your webhooks.","status":"GA","models":[{"name":"AI Agent Runtime","status":"GA","notes":"$0.16/min including real-time STT, LLM and standard TTS; orchestration, barge-in, memory, tool calls and retrieval included. Voice transport billed separately."}],"transports":["PSTN","SIP","WebRTC"],"audio":{"input":"Phone/SIP/WebRTC","output":"Standard TTS included; ElevenLabs/premium TTS extra"},"languages":"Configurable via ai.languages or ai.multilingual.","latency":"No verified number.","features":["telephony","function calling (SWAIG)","post-call summary webhook (post_prompt_url)","barge-in","memory","retrieval","pronunciation hints"],"pricing":{"model":"per-minute","items":[{"what":"AI Agent Runtime","price":"$0.16","unit":"per minute","notes":"Includes STT, LLM and standard TTS."},{"what":"Local inbound / outbound","price":"$0.0066 / $0.008","unit":"per minute","notes":"SIP and WebRTC $0.003/min."},{"what":"Local number","price":"$0.50","unit":"per month","notes":"Toll-free $0.80/month; port-out $5/number."},{"what":"ElevenLabs TTS","price":"$0.000297","unit":"per character","notes":"High quality $0.000594/char; page is inconsistent on how premium TTS is billed for agents."}],"est_per_minute_usd":{"low":0.163,"high":0.168,"basis":"SignalWire's own all-in estimate for standard SIP/PSTN usage with standard TTS. Premium/ElevenLabs voices add more."},"free_tier":"Not stated on pricing page.","source":"https://signalwire.com/pricing"},"limits":["Not published in sources reviewed"],"regions":"Not verified.","setup":{"steps":["Create a SignalWire space and buy a number.","Create a SWML script (dashboard or hosted URL) with answer + ai.","Point the number's call handler at the SWML; add SWAIG functions that call your webhooks; set post_prompt_url for call summaries."],"endpoint":"SWML served from your URL or the dashboard","auth":"Project ID + API token","snippet_lang":"javascript","snippet":"// Serve a minimal SWML AI agent (JSON form of SWML) from Express\nimport express from \"express\";\nconst app = express();\napp.all(\"/swml\", (req, res) => res.json({\n  version: \"1.0.0\",\n  sections: {\n    main: [\n      { answer: {} },\n      { ai: {\n          prompt: { text: \"You are a customer service agent for Acme. Say you are an AI. Be brief.\" },\n          post_prompt: { text: \"Summarize the call and any actions taken.\" },\n          post_prompt_url: \"https://example.com/post-prompt-hook\"\n      } }\n    ]\n  }\n}));\napp.listen(3000);"},"warnings":[{"severity":"medium","title":"Premium voices cost extra","detail":"Standard TTS is included; ElevenLabs is billed per character and the page is inconsistent about how it applies to agents. Price your exact voice."},{"severity":"medium","title":"Higher headline than DIY","detail":"$0.16/min bundled is roughly 2x a DIY Twilio/Plivo + cheap model stack, but includes the LLM."},{"severity":"low","title":"SWML learning curve","detail":"Agent behaviour lives in SWML/SWAIG rather than code, which is fast to start but different from Python/JS frameworks."},{"severity":"low","title":"LLM choice","detail":"The bundled LLM is chosen by SignalWire; confirm which models and whether you can bring your own."}],"best_for":"Phone agents where one predictable bundled rate including the LLM is preferred.","open_source":false,"self_hostable":false,"compliance":"Not verified in this research.","docs":[{"label":"Pricing","url":"https://signalwire.com/pricing"},{"label":"SWML ai method","url":"https://signalwire.com/docs/swml/reference/calling/ai"}],"sources":["https://signalwire.com/pricing","https://signalwire.com/docs/swml/reference/calling/ai"],"confidence":"medium","unverified":"Which LLM is bundled; free trial; compliance.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.16,"all_in":true,"phone_numbers":true,"byo_llm":null,"byo_keys":null,"no_code_builder":null,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"$0.16/min includes STT, LLM and standard TTS; voice transport and premium voices extra. Bundled LLM not named.","_added":[]}},{"id":"openai-agents-sdk-voice","name":"OpenAI Agents SDK (voice / realtime agents)","vendor":"OpenAI","category":"framework","summary":"Free open-source SDK (Python and TypeScript) with RealtimeAgent/RealtimeSession wrappers over the OpenAI Realtime API (browser WebRTC with ephemeral keys, server WebSocket, SIP), and a Python VoicePipeline for chained STT -> agent -> TTS.","status":"GA","models":[{"name":"RealtimeAgent / RealtimeSession (JS: @openai/agents/realtime)","status":"GA","notes":"Handles interruptions, tools and handoffs between realtime agents."},{"name":"VoicePipeline (Python, openai-agents[voice])","status":"GA","notes":"Chained STT -> agent workflow -> TTS. The JS README lists a voice pipeline as a future item."}],"transports":["WebRTC","WebSocket","SIP"],"audio":{"input":"Mic via WebRTC in browser; PCM/G.711 over WebSocket","output":"Same"},"languages":"Per OpenAI realtime/transcription models.","latency":"Model-dependent.","features":["tools","handoffs","guardrails","interruptions","tracing","SIP"],"pricing":{"model":"free/open","items":[{"what":"SDK","price":"$0","unit":"license","notes":"You pay OpenAI Realtime/STT/TTS model usage (see model segment)."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Cost is the OpenAI model bill, priced in the realtime models segment."},"free_tier":"Open source.","source":"https://www.npmjs.com/package/@openai/agents"},"limits":["OpenAI Realtime rate limits apply"],"regions":"OpenAI API regions.","setup":{"steps":["npm i @openai/agents (or pip install 'openai-agents[voice]').","Server: mint a short-lived client key for the Realtime API; never ship your real API key to the browser.","Browser: create a RealtimeAgent and RealtimeSession and connect with the ephemeral key (WebRTC, mic handled for you)."],"endpoint":"OpenAI Realtime API","auth":"Ephemeral client secret in browser; API key server-side","snippet_lang":"javascript","snippet":"// Browser voice agent with the OpenAI Agents SDK (TypeScript/JS)\nimport { RealtimeAgent, RealtimeSession } from \"@openai/agents/realtime\";\n\nconst agent = new RealtimeAgent({\n  name: \"Assistant\",\n  instructions: \"You are a friendly voice assistant. Keep answers short.\",\n});\nconst session = new RealtimeSession(agent);\n\n// ephemeralKey comes from your server (OpenAI client secret endpoint)\nconst { ephemeralKey } = await fetch(\"/api/realtime-token\").then(r => r.json());\nawait session.connect({ apiKey: ephemeralKey }); // WebRTC + mic in the browser"},"warnings":[{"severity":"medium","title":"Ties you to OpenAI realtime pricing","detail":"The SDK is free but speech-to-speech realtime minutes are among the more expensive options; long sessions also grow context cost."},{"severity":"medium","title":"Never expose the API key","detail":"Browser sessions must use short-lived client secrets minted server-side."},{"severity":"low","title":"Feature parity differs by language","detail":"VoicePipeline is documented for Python; check JS docs before relying on it."},{"severity":"low","title":"No telephony numbers included","detail":"SIP support exists, but you still need a carrier (Twilio, Telnyx, etc.)."}],"best_for":"Teams standardised on OpenAI models who want agent tooling/handoffs with realtime voice.","open_source":true,"self_hostable":true,"compliance":"Per OpenAI API terms (not covered here).","docs":[{"label":"npm @openai/agents","url":"https://www.npmjs.com/package/@openai/agents"},{"label":"Python SDK","url":"https://openai.github.io/openai-agents-python/"},{"label":"JS SDK","url":"https://openai.github.io/openai-agents-js/"}],"sources":["https://www.npmjs.com/package/@openai/agents","https://cdn.jsdelivr.net/gh/openai/openai-agents-python@main/README.md"],"confidence":"medium","unverified":"Current default realtime model name; whether JS has a chained VoicePipeline yet.","cat":"platforms","kind":"framework","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":true,"open_source":true,"self_hostable":true,"platform_fee_per_min":0,"all_in":false,"phone_numbers":false,"byo_llm":null,"byo_keys":null,"no_code_builder":null,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Free SDK; cost is OpenAI model usage. SIP supported but you need a carrier for phone numbers.","_added":[]}},{"id":"vercel-ai-sdk-realtime","name":"Vercel AI SDK realtime (via AI Gateway)","vendor":"Vercel","category":"framework","summary":"Beta realtime voice support in the AI SDK through Vercel AI Gateway: gateway.experimental_realtime mints connection tokens server-side and a useRealtime React hook handles mic capture and playback for speech-to-speech models.","status":"Beta","models":[{"name":"gateway.experimental_realtime + useRealtime","status":"Beta","notes":"Changelog (29 Jun 2026) says available via AI SDK 7; docs say install canary packages. Examples use openai/gpt-realtime-2 and xai/grok-voice-think-fast-1.0."}],"transports":["WebSocket","WebRTC"],"audio":{"input":"Browser mic","output":"Browser playback"},"languages":"Per model.","latency":"Model-dependent.","features":["one gateway key for multiple realtime model vendors","React hook","batch STT/TTS via AI Gateway"],"pricing":{"model":"usage","items":[{"what":"AI SDK","price":"$0","unit":"license","notes":"AI Gateway passes through model usage; see Vercel AI Gateway pricing and model segment."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Equals the underlying realtime model cost; see models segment."},"free_tier":"SDK open source; AI Gateway has its own terms.","source":"https://vercel.com/docs/ai-gateway/modalities/realtime"},"limits":["Beta / canary API surface","Some models are speech-to-speech only (no transcription)"],"regions":"Vercel AI Gateway.","setup":{"steps":["pnpm add ai@canary @ai-sdk/gateway@canary @ai-sdk/react@canary (or AI SDK 7 per changelog).","Create a server route that uses gateway.experimental_realtime getToken to mint a connection.","Use the useRealtime hook in a React client for mic and playback."],"endpoint":"Vercel AI Gateway","auth":"AI_GATEWAY_API_KEY or Vercel OIDC","snippet_lang":"javascript","snippet":"// Vercel AI SDK realtime is Beta and its signatures are changing; follow\n// https://vercel.com/docs/ai-gateway/modalities/realtime for current code.\n// Install: pnpm add ai@canary @ai-sdk/gateway@canary @ai-sdk/react@canary\n// Server: gateway.experimental_realtime -> getToken() for a model such as \"openai/gpt-realtime-2\"\n// Client: const rt = useRealtime({ ...token endpoint... }) -> start/stop mic + playback"},"warnings":[{"severity":"medium","title":"Canary/Beta API","detail":"Exact function signatures were not stable enough to publish a full snippet; expect breaking changes."},{"severity":"low","title":"Version confusion","detail":"Docs say canary releases; changelog says AI SDK 7. Use whichever the docs page you follow specifies."},{"severity":"low","title":"Speech-to-speech only models","detail":"Some gateway realtime models do not support transcription, so you cannot get user transcripts from them."}],"best_for":"Next.js/React apps already using the AI SDK and AI Gateway that want a quick browser voice mode.","open_source":true,"self_hostable":false,"compliance":"Not covered.","docs":[{"label":"AI Gateway realtime","url":"https://vercel.com/docs/ai-gateway/modalities/realtime"},{"label":"Changelog","url":"https://vercel.com/changelog/realtime-voice-speech-and-transcription-now-supported-on-ai-gateway"}],"sources":["https://vercel.com/docs/ai-gateway/modalities/realtime","https://vercel.com/changelog/realtime-voice-speech-and-transcription-now-supported-on-ai-gateway"],"confidence":"low","unverified":"Exact API signatures; gateway markup on realtime models.","cat":"platforms","kind":"framework","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":true,"self_hostable":false,"platform_fee_per_min":0,"all_in":false,"phone_numbers":false,"byo_llm":null,"byo_keys":null,"no_code_builder":null,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Beta/canary API. SDK is free; AI Gateway passes through model cost (markup on realtime models unverified).","_added":[]}},{"id":"layercode","name":"Layercode","vendor":"Layercode, Inc.","category":"platform","summary":"Voice-agent pipeline platform that connected your own agent backend (webhook) to a hosted STT/TTS voice layer. As of 2026-10-10 layercode.com and docs.layercode.com redirect to toyo.ai, a different AI-agent product from the same company.","status":"Deprecated","models":[{"name":"Layercode voice pipeline","status":"Deprecated (unclear)","notes":"No public pricing or docs reachable; footer of toyo.ai says 'Toyo is a trademark of Layercode, Inc.'"}],"transports":["WebSocket","WebRTC"],"audio":{"input":"Unknown (site unavailable)","output":"Unknown"},"languages":"Unknown.","latency":"Unknown.","features":[],"pricing":{"model":"usage","items":[],"est_per_minute_usd":{"low":null,"high":null,"basis":"No pricing reachable."},"free_tier":"Unknown.","source":"https://toyo.ai/"},"limits":[],"regions":"Unknown.","setup":{"steps":["Do not start new projects on Layercode until the vendor confirms the voice product is still offered."],"endpoint":"","auth":"","snippet_lang":"javascript","snippet":"// Layercode docs now redirect to toyo.ai; no current snippet available."},"warnings":[{"severity":"high","title":"Product site redirects elsewhere","detail":"layercode.com/pricing and docs.layercode.com 301-redirect to toyo.ai, which does not mention a voice API. Existing users should confirm service continuity and plan a migration path."},{"severity":"medium","title":"No shutdown notice found","detail":"No official deprecation post was found, so the status is inferred from the redirects."}],"best_for":"Not recommended for new builds until status is clarified.","open_source":false,"self_hostable":false,"compliance":"Unknown.","docs":[{"label":"Toyo (redirect target)","url":"https://toyo.ai/"}],"sources":["https://layercode.com/pricing","https://docs.layercode.com/","https://toyo.ai/"],"confidence":"medium","unverified":"Whether the Layercode voice API still serves existing customers.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"platform_fee_per_min":null,"all_in":null,"phone_numbers":false,"byo_llm":true,"byo_keys":null,"no_code_builder":null,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Status unclear: layercode.com now redirects to toyo.ai; no pricing or docs reachable on 2026-10-10.","_added":[]}},{"id":"vapi","name":"Vapi","vendor":"Vapi","category":"platform","summary":"Developer-first hosted voice-agent platform: configure transcriber, LLM, voice and tools per assistant (dashboard or API), connect phone numbers or web/SDK clients; Vapi charges a hosting fee and passes model and carrier costs through.","status":"GA","models":[{"name":"No Success Package (pay as you go)","status":"GA","notes":"$0 to start, $0.05/min hosting, $5 free credits, 1 phone number, 4 concurrent calls, 14-day retention."},{"name":"Core","status":"GA","notes":"$29/month: 10 concurrent calls, 5 numbers, 30-day retention, zero data retention option."},{"name":"Pro","status":"GA","notes":"10% of hosting fee with $999/month minimum: 30 concurrent calls, RBAC, 99% uptime SLA, 180-day retention."},{"name":"Premier","status":"GA","notes":"Custom: SSO, dedicated SIP, 99.9% SLA."}],"transports":["WebRTC","WebSocket","SIP","PSTN"],"audio":{"input":"Phone (Vapi numbers, Twilio, Vonage, Telnyx, SIP), Daily WebRTC web SDK, WebSocket","output":"Same"},"languages":"Depends on transcriber/voice chosen.","latency":"No verified number.","features":["turn detection / endpointing","interruptions","tools/function calling","squads (multi-assistant)","telephony","multi-provider plug-ins","recording","web SDKs","HIPAA mode (hipaaEnabled)","zero data retention (Core+)"],"pricing":{"model":"per-minute","items":[{"what":"Vapi hosting","price":"$0.05","unit":"per minute","notes":"Platform fee only."},{"what":"Transcriber (Deepgram)","price":"$0.0095-$0.0099","unit":"per minute","notes":"Passed through at provider cost (no markup per Vapi), or bring your own key."},{"what":"LLM (OpenAI)","price":"$0.0077-$0.0452","unit":"per minute","notes":"Pass-through."},{"what":"Voice (ElevenLabs)","price":"$0.0146-$0.0238","unit":"per minute","notes":"Pass-through."},{"what":"Telephony","price":"Twilio $0.008 in / $0.014 out; Vonage $0.00814; Telnyx $0.0055","unit":"per minute","notes":"Charged by the carrier, not Vapi. Vapi SIP/WebSocket/Daily WebRTC transport free."},{"what":"HIPAA add-on","price":"$2,000","unit":"per month","notes":"Any package."},{"what":"Extra concurrency","price":"$10","unit":"per line per month","notes":""}],"est_per_minute_usd":{"low":0.09,"high":0.15,"basis":"Vapi's estimator: 1,000 minutes = $82-$129 including $50 hosting (so $0.082-$0.129/min before carrier), plus $0.008-$0.014 Twilio. Premium LLMs or voices push higher."},"free_tier":"$5 free credits and 1 phone number.","source":"https://www.vapi.ai/pricing"},"limits":["Concurrency: 4 (free) / 10 (Core) / 30 (Pro), $10 per extra line/month","Phone numbers: 1 / 5 / 10","Log retention 14 / 30 / 180 days"],"regions":"Not verified.","setup":{"steps":["Sign up, create an assistant in the dashboard (transcriber, model, voice, first message, tools).","Web: npm i @vapi-ai/web and start a call with your public key and assistant ID.","Phone: buy/import a number and attach the assistant, or create outbound calls with @vapi-ai/server-sdk (calls.create).","Add server URL webhooks for tool calls and end-of-call reports."],"endpoint":"https://api.vapi.ai","auth":"Public key in browser; private API key server-side (Bearer token)","snippet_lang":"javascript","snippet":"// Browser: talk to an assistant\nimport Vapi from \"@vapi-ai/web\";\nconst vapi = new Vapi(import.meta.env.VITE_VAPI_PUBLIC_KEY);\nvapi.on(\"call-start\", () => console.log(\"connected\"));\nvapi.on(\"message\", (m) => console.log(m));      // transcripts, tool calls, status\nvapi.on(\"call-end\", () => console.log(\"ended\"));\ndocument.querySelector(\"#talk\").onclick = () => vapi.start(\"YOUR_ASSISTANT_ID\");\n\n// Server: place an outbound phone call\nimport { VapiClient } from \"@vapi-ai/server-sdk\";\nconst client = new VapiClient({ token: process.env.VAPI_API_KEY });\nawait client.calls.create({\n  phoneNumberId: \"YOUR_PHONE_NUMBER_ID\",\n  customer: { number: \"+15555550123\" },\n  assistantId: \"YOUR_ASSISTANT_ID\",\n});"},"warnings":[{"severity":"high","title":"$0.05/min is about half the real cost","detail":"Vapi's own estimator lands at $0.082-$0.129/min before telephony. Always model the full stack."},{"severity":"high","title":"HIPAA is a $2,000/month add-on","detail":"Turning on hipaaEnabled is not enough for PHI; budget the add-on and BAAs with each provider."},{"severity":"medium","title":"Free tier allows only 4 concurrent calls","detail":"Outbound campaigns need Core ($29) or extra lines at $10/line/month."},{"severity":"medium","title":"Pro is a $999/month minimum","detail":"The 10% of hosting fee pricing only makes sense at high volume."},{"severity":"low","title":"Free numbers are US-only and limited","detail":"One included number on the free package; international numbers come from your own carrier."}],"best_for":"Developers who want a hosted pipeline with full provider choice and an API-first workflow.","open_source":false,"self_hostable":false,"compliance":"SOC 2 Type II per third-party trust listing (security.vapi.ai trust center); HIPAA via $2,000/mo add-on; GDPR page in docs.","docs":[{"label":"Pricing","url":"https://www.vapi.ai/pricing"},{"label":"Web quickstart","url":"https://docs.vapi.ai/quickstart/web"},{"label":"HIPAA","url":"https://docs.vapi.ai/security-and-privacy/hipaa"}],"sources":["https://www.vapi.ai/pricing","https://docs.vapi.ai/quickstart/web","https://docs.vapi.ai/security-and-privacy/hipaa","https://trustlists.org/company/vapi/"],"confidence":"high","unverified":"SOC 2 report type confirmed only via third-party listing; regions.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","short":"Vapi","facts":{"latency_ms":500,"languages":null,"max_session_min":null,"concurrency":10,"free_tier":true,"free_credit_usd":5,"entry_plan_usd_month":29,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":true,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.05,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":true,"no_code_builder":null,"recording":true,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Latency is a vendor blog claim (sub-500 ms). Free tier 4 concurrent, Core $29 gives 10. Real all-in cost $0.08-$0.13/min before telephony. HIPAA is a $2,000/month add-on. SOC 2 confirmed via trust center listing.","_added":["latency_ms: https://vapi.ai/blog/speech-latency"]}},{"id":"retell-ai","name":"Retell AI","vendor":"Retell AI","category":"platform","summary":"Hosted voice-agent platform for phone and web calls with a component price list: voice infrastructure fee, chosen LLM per minute, chosen TTS per minute, and telephony; strong on contact-center features like batch calls, QA and branded caller ID.","status":"GA","models":[{"name":"Pay as you go","status":"GA","notes":"$10 free credits; 20 concurrent calls included; $0.07-$0.31/min typical range per Retell."},{"name":"Enterprise","status":"GA","notes":"Custom pricing, concurrency from 50+."}],"transports":["WebRTC","SIP","PSTN","WebSocket"],"audio":{"input":"Phone (Retell numbers, Twilio, Telnyx, custom SIP), web call SDK, custom LLM WebSocket","output":"Same"},"languages":"Depends on chosen voice/LLM; not verified.","latency":"No verified number.","features":["turn detection","interruptions","conversation flow builder","knowledge base","advanced denoising","PII removal","safety guardrails","AI quality assurance","batch calls","branded caller ID","SMS","custom LLM via WebSocket","telephony"],"pricing":{"model":"per-minute","items":[{"what":"Retell voice infrastructure","price":"$0.055","unit":"per minute","notes":"Includes STT and the base voice engine; LLM and TTS selected separately."},{"what":"LLM","price":"GPT 4.1 nano $0.0032 ... GPT 4.1 mini $0.0128 ... GPT 5.4 $0.08 ... GPT 5.5 $0.16","unit":"per minute","notes":"Fast tier roughly doubles some rates. Bring your own LLM via custom LLM WebSocket."},{"what":"TTS","price":"$0.015 (Retell/Minimax/Fish/Cartesia/OpenAI/Inworld); ElevenLabs Flash $0.04; Multilingual v2 $0.06; v3 $0.10","unit":"per minute","notes":""},{"what":"Telephony","price":"$0.015","unit":"per minute","notes":"Varies by country; no charge for SIP trunking/custom telephony."},{"what":"Extra concurrency","price":"$8","unit":"per concurrent call per month","notes":"20 included."},{"what":"Add-ons","price":"+$0.005 each (KB, denoising, guardrails); PII +$0.01; QA $0.10 after 100 min","unit":"per minute","notes":"Batch call +$0.005/dial; branded call +$0.10/outbound call."},{"what":"Phone number","price":"$2.00","unit":"per month","notes":"Custom numbers free."}],"est_per_minute_usd":{"low":0.09,"high":0.31,"basis":"Low: $0.055 infra + GPT 4.1 mini $0.0128 + $0.015 TTS + $0.015 telephony = ~$0.098 (cheaper with nano and BYO telephony). High: Retell's own top of range, e.g. GPT 5.x + ElevenLabs v3 + add-ons."},"free_tier":"$10 free credits.","source":"https://www.retellai.com/pricing"},"limits":["20 concurrent calls included, $8/month per extra","Retell-purchased numbers can only dial US destinations via API create-phone-call"],"regions":"Not verified.","setup":{"steps":["Create an agent in the dashboard (single prompt or conversation flow), pick LLM and voice.","Buy or import a number and bind the agent for inbound; or call create-phone-call for outbound.","For web, create a web call server-side and start it with the Retell web SDK.","Sign the BAA/DPA click-through before handling PHI."],"endpoint":"POST https://api.retellai.com/v2/create-phone-call","auth":"Authorization: Bearer <RETELL_API_KEY>","snippet_lang":"javascript","snippet":"// Outbound call with Retell\nconst res = await fetch(\"https://api.retellai.com/v2/create-phone-call\", {\n  method: \"POST\",\n  headers: {\n    Authorization: \"Bearer \" + process.env.RETELL_API_KEY,\n    \"Content-Type\": \"application/json\",\n  },\n  body: JSON.stringify({\n    from_number: \"+14155550100\",           // a number you own in Retell\n    to_number: \"+14155550123\",\n    override_agent_id: \"agent_xxx\",\n    retell_llm_dynamic_variables: { customer_name: \"Sam\" },\n    idempotency_key: \"appt-reminder-7781\",  // safe retries for 1 hour\n    honor_internal_dnc: true,\n  }),\n});\nconsole.log(res.status, await res.json()); // 201 -> call_id, call_status"},"warnings":[{"severity":"high","title":"Silence and hold are billed","detail":"Billing runs per second for the whole call including silence and hold; set end-call-on-silence and max duration."},{"severity":"medium","title":"LLM choice dominates cost","detail":"GPT 5.5 at $0.16/min alone is 3x the infra fee; test cheaper models (4.1 mini, Gemini Flash) first."},{"severity":"medium","title":"Add-ons stack up","detail":"Knowledge base, denoising, guardrails (+$0.005 each), PII removal (+$0.01) and branded calls ($0.10/call) can add 20-40% to the per-minute rate."},{"severity":"low","title":"Transfers continue telephony billing","detail":"After a transfer only the telephony fee continues, but it does continue."},{"severity":"low","title":"International outbound needs your own numbers","detail":"Retell-purchased numbers support US destinations only for API outbound calls."}],"best_for":"Teams running production phone agents (inbound support, outbound reminders) that want contact-center features without building them.","open_source":false,"self_hostable":false,"compliance":"SOC 2 Type 1 and Type 2, HIPAA (self-serve BAA at click-agreements.retellai.com), GDPR DPA with SCCs, per Retell docs.","docs":[{"label":"Pricing","url":"https://www.retellai.com/pricing"},{"label":"Create phone call API","url":"https://docs.retellai.com/api-references/create-phone-call"},{"label":"Compliance","url":"https://docs.retellai.com/general/compliance"}],"sources":["https://www.retellai.com/pricing","https://docs.retellai.com/api-references/create-phone-call","https://docs.retellai.com/general/compliance"],"confidence":"high","unverified":"Exactly what is included in the $0.055 voice infra (STT assumed); languages list.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","short":"Retell AI","facts":{"latency_ms":600,"languages":null,"max_session_min":null,"concurrency":20,"free_tier":true,"free_credit_usd":10,"entry_plan_usd_month":0,"hipaa":true,"soc2":true,"gdpr_eu":true,"webrtc":true,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.055,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":null,"no_code_builder":true,"recording":null,"noise_cancellation":true,"turn_detection_model":null,"_notes":"Latency is vendor claim (as low as 600 ms). $0.055/min voice infra includes STT; LLM and TTS priced per minute on top. 20 concurrent included, $8/month per extra. Typical $0.07-$0.31/min all-in.","_added":["latency_ms: https://docs.retellai.com/reliability/check-actual-latency.md"]}},{"id":"bland-ai","name":"Bland AI","vendor":"Bland","category":"platform","summary":"Phone-first AI calling platform with one bundled per-minute rate covering STT, LLM and TTS, conversational 'pathways', and plan-based daily/hourly call caps.","status":"GA","models":[{"name":"Start","status":"GA","notes":"Free plan: $0.14/min connected, $0.05/min transfer, 100 calls/day, 10 concurrent, 1 voice clone. Includes 2 credits + an inbound number."},{"name":"Build","status":"GA","notes":"$299/month: $0.12/min, $0.04/min transfer, 2,000 calls/day, 1,000/hour, 50 concurrent."},{"name":"Scale","status":"GA","notes":"$499/month: $0.11/min, $0.03/min transfer, 5,000 calls/day, 100 concurrent (listed in billing docs, not on the marketing pricing page)."},{"name":"Agent Phone","status":"GA","notes":"$29.99/month: 1 concurrent call, unlimited US/Canada voice minutes up to 1,000 call minutes/day."},{"name":"Enterprise","status":"GA","notes":"Custom; SMS and web chat included."}],"transports":["PSTN","SIP","WebRTC"],"audio":{"input":"Phone (Bland numbers or BYO Twilio/SIP), web widget","output":"Same"},"languages":"Not verified.","latency":"No verified number.","features":["pathways (conversation graphs)","voice cloning","transfers","knowledge bases","SMS (Enterprise)","recording","BYO telephony","telephony"],"pricing":{"model":"per-minute","items":[{"what":"Connected minute (Start / Build / Scale)","price":"$0.14 / $0.12 / $0.11","unit":"per minute","notes":"Includes LLM, STT and TTS; no token charges. Rates effective 5 Dec 2025 (previous standard $0.09)."},{"what":"Transfer time","price":"$0.05 / $0.04 / $0.03","unit":"per minute","notes":"Free with BYO telephony."},{"what":"Outbound minimum","price":"$0.015","unit":"per call","notes":"On Bland telephony; absorbed once a call exceeds ~10 seconds at the per-minute rate."},{"what":"Failed call (Bland telephony)","price":"$0.015","unit":"per call","notes":""},{"what":"SMS","price":"$0.02","unit":"per message","notes":""},{"what":"Telephony","price":"Pass-through (Bland's Twilio) or your carrier","unit":"per minute","notes":"Marketing page says telephony is billed separately at pass-through cost; billing docs list SIP termination $0.004/min as absorbed by Bland."},{"what":"Plans","price":"$0 / $299 / $499","unit":"per month","notes":"Start / Build / Scale."}],"est_per_minute_usd":{"low":0.11,"high":0.16,"basis":"$0.11-$0.14 bundled AI rate plus pass-through carrier minutes (roughly $0.004-$0.014). Short or failed outbound attempts cost $0.015 each regardless."},"free_tier":"Start plan: free, no card, 2 credits + an inbound number.","source":"https://docs.bland.ai/platform/billing"},"limits":["Daily caps 100 / 2,000 / 5,000 calls; hourly caps 100 / 1,000 / 1,000","Concurrency 10 / 50 / 100","International calls need at least $5 credit purchase or auto-recharge","Agent Phone numbers call US/Canada only"],"regions":"Not verified.","setup":{"steps":["Sign up (Start plan) and get an API key.","Write a task prompt or build a pathway in the dashboard.","POST /v1/calls with phone_number and task (or pathway_id); poll or receive webhooks for transcripts and recording_url."],"endpoint":"POST https://api.bland.ai/v1/calls","auth":"authorization: <API key> (Bearer prefix optional)","snippet_lang":"javascript","snippet":"// Place a Bland AI call\nconst res = await fetch(\"https://api.bland.ai/v1/calls\", {\n  method: \"POST\",\n  headers: { authorization: process.env.BLAND_API_KEY, \"Content-Type\": \"application/json\" },\n  body: JSON.stringify({\n    phone_number: \"+14155550123\",\n    task: \"You are an AI assistant from Acme Dental confirming tomorrow's 3pm appointment. Say you are an AI. If they want to reschedule, offer Thursday or Friday.\",\n    first_sentence: \"Hi, this is Acme Dental's AI assistant calling about your appointment.\",\n    voice: \"Karen\",\n    max_duration: 5,      // minutes; default is 30\n    record: true,\n  }),\n});\nconsole.log(await res.json());"},"warnings":[{"severity":"high","title":"Prices rose in Dec 2025","detail":"Older blogs quote $0.09/min; current rates are $0.11-$0.14/min depending on plan, plus transfer minutes."},{"severity":"medium","title":"$0.015 per outbound attempt","detail":"No-answers, busy and failed calls on Bland telephony each cost $0.015; a 10,000-number campaign with 70% no-answer adds about $105 before any talk time."},{"severity":"medium","title":"Hard daily and hourly caps","detail":"Start allows 100 calls/day; Build 2,000/day and 1,000/hour. Plan campaigns around caps, not just concurrency."},{"severity":"medium","title":"Default max_duration is 30 minutes","detail":"A stuck call (voicemail loop, hold music) can run up to 30 minutes at the full rate; lower it."},{"severity":"low","title":"Telephony wording is inconsistent","detail":"Marketing page says pass-through telephony billed separately; billing docs show SIP termination absorbed. Check an invoice early."}],"best_for":"Outbound and inbound phone automation where a single bundled rate is easier than managing providers.","open_source":false,"self_hostable":false,"compliance":"Not verified in this research (ask Bland for SOC 2/HIPAA terms).","docs":[{"label":"Pricing","url":"https://www.bland.ai/pricing"},{"label":"Billing docs","url":"https://docs.bland.ai/platform/billing"},{"label":"Send call API","url":"https://docs.bland.ai/api-v1/post/calls"}],"sources":["https://www.bland.ai/pricing","https://docs.bland.ai/platform/billing","https://docs.bland.ai/api-v1/post/calls","https://docs.bland.ai/changelog/06_09_2025"],"confidence":"high","unverified":"Exact telephony pass-through rates; compliance attestations; phone number price.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":50,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":299,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.14,"all_in":true,"phone_numbers":true,"byo_llm":null,"byo_keys":null,"no_code_builder":true,"recording":true,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Bundled rate $0.14 (free Start), $0.12 (Build $299), $0.11 (Scale $499); carrier minutes and $0.015 per outbound attempt extra. Concurrency 10/50/100. A $29.99 Agent Phone plan (1 concurrent) also exists. Default max_duration 30 min (configurable).","_added":[]}},{"id":"synthflow","name":"Synthflow","vendor":"Synthflow AI","category":"platform","summary":"No-code voice-agent builder aimed at businesses and agencies; as of October 2026 the official pricing page shows only an Enterprise plan from $30,000 per year, with no public per-minute rate.","status":"GA","models":[{"name":"Enterprise","status":"GA","notes":"Contracts start at $30,000/year; priced on volume, concurrency, telephony, integrations, security."}],"transports":["PSTN","SIP"],"audio":{"input":"Phone (Synthflow native telephony, SIP trunking)","output":"Same"},"languages":"Not verified.","latency":"No verified number.","features":["no-code builder","telephony","CRM/calendar integrations","routing and handoffs","webhooks/API"],"pricing":{"model":"subscription","items":[{"what":"Enterprise contract","price":"from $30,000","unit":"per year","notes":"Per-minute rates, model inclusion and concurrency not published."}],"est_per_minute_usd":{"low":0.13,"high":0.24,"basis":"Third-party estimates only (CloudTalk, Layer3 Labs: ~$0.08-$0.09 base plus LLM/telephony, effective $0.13-$0.24/min). Not verified on synthflow.ai; treat as indicative."},"free_tier":"None published.","source":"https://www.synthflow.ai/pricing"},"limits":["Concurrency defined per contract"],"regions":"Not verified.","setup":{"steps":["Contact sales; enterprise onboarding includes implementation and testing.","Build agents in the no-code studio, attach numbers or SIP trunks, connect CRM/calendar."],"endpoint":"Synthflow API (enterprise)","auth":"API key","snippet_lang":"javascript","snippet":"// Synthflow is configured in its no-code studio; API access is part of enterprise onboarding.\n// No public self-serve quickstart was available on 2026-10-10."},"warnings":[{"severity":"high","title":"No self-serve pricing anymore","detail":"The public pricing page lists only Enterprise from $30,000/year. Blog posts quoting $0.08-$0.09/min pay-as-you-go may be outdated for new customers."},{"severity":"medium","title":"Per-minute add-ons reported","detail":"Third parties report add-ons (routing, low-latency edge) at extra per-minute cost that can push effective rates above $0.20/min."},{"severity":"low","title":"HIPAA only as a badge","detail":"HIPAA appears as a footer trust badge; confirm BAA terms in the contract."}],"best_for":"Businesses that want a vendor-managed, no-code rollout and can commit to an annual contract.","open_source":false,"self_hostable":false,"compliance":"HIPAA badge on site; MSA/DPA support per pricing page; details not verified.","docs":[{"label":"Pricing","url":"https://www.synthflow.ai/pricing"}],"sources":["https://www.synthflow.ai/pricing","https://www.cloudtalk.io/blog/synthflow-pricing/","https://www.layer3labs.io/guides/synthflow-pricing","https://usagepricing.com/blueprint/synthflow"],"confidence":"low","unverified":"Per-minute rates, concurrency, included models, whether legacy pay-as-you-go accounts still exist.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":2500,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":null,"all_in":null,"phone_numbers":true,"byo_llm":null,"byo_keys":null,"no_code_builder":true,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Only Enterprise listed, from $30,000/year (about $2,500/month). No public per-minute rate; third-party estimates $0.13-$0.24/min. HIPAA shown only as a site badge.","_added":[]}},{"id":"millis-ai","name":"Millis AI","vendor":"Millis AI","category":"platform","summary":"Low-latency voice-agent platform with a $0.02/min platform fee on top of separately priced STT, LLM and character-metered TTS; bring your own LLM endpoint at no LLM charge.","status":"GA","models":[{"name":"Pay as you go","status":"GA","notes":"$0.02/min platform + components; no subscription tier published."}],"transports":["WebRTC","WebSocket","PSTN","SIP"],"audio":{"input":"Web SDK and phone (via your Twilio/Plivo etc.)","output":"Same"},"languages":"Not verified.","latency":"Vendor markets very low latency; no verified number.","features":["multi-provider plug-ins","custom LLM endpoint","telephony integrations","'Choose by Millis' lowest-latency model selection"],"pricing":{"model":"per-minute","items":[{"what":"Platform fee","price":"$0.02","unit":"per minute","notes":"On top of LLM, TTS and STT."},{"what":"STT","price":"$0.0043","unit":"per minute","notes":""},{"what":"LLM","price":"GPT-4o ~$0.004, GPT-3.5 Turbo ~$0.0004, Llama-3 ~$0.00018","unit":"per minute (estimate)","notes":"Custom LLM endpoint: no Millis charge. Model list on the docs page looks dated."},{"what":"TTS","price":"ElevenLabs $0.10, Cartesia $0.0392, OpenAI/Deepgram $0.015, Rime $0.075","unit":"per 1,000 characters","notes":"Roughly $0.0075-$0.05/min per Millis' own conversion."}],"est_per_minute_usd":{"low":0.04,"high":0.09,"basis":"$0.02 + $0.0043 STT + ~$0.004 LLM + $0.0075-$0.05 TTS + your carrier minutes (~$0.005-$0.014). Third-party worked example: ~$0.066/min with GPT-4o + ElevenLabs."},"free_tier":"Not published.","source":"https://docs.millis.ai/pricing"},"limits":["Concurrency not published"],"regions":"Not verified.","setup":{"steps":["Sign up and create an agent (prompt, voice, LLM) in the dashboard.","Use the web SDK for browser calls or connect a Twilio/Plivo number for phone.","Optionally point the agent at your own LLM endpoint."],"endpoint":"Millis API","auth":"Public key (web SDK) / private API key","snippet_lang":"javascript","snippet":"// Millis agents are created in the dashboard; see docs.millis.ai for the current web SDK snippet.\n// Snippet not reproduced here because the SDK surface could not be verified on 2026-10-10."},"warnings":[{"severity":"medium","title":"TTS billed per character","detail":"Cost depends on how much the agent talks; a chatty agent with ElevenLabs costs several times the headline."},{"severity":"medium","title":"Pricing page looks stale","detail":"Docs still list GPT-4 Turbo, GPT-3.5 and PlayHT; confirm current models and rates in the dashboard."},{"severity":"low","title":"Telephony not priced","detail":"Phone minutes come from your own carrier account."},{"severity":"low","title":"Limited public compliance info","detail":"No SOC 2/HIPAA statements were verified."}],"best_for":"Cost-sensitive builders who want a cheap orchestration fee and their own LLM.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Pricing","url":"https://docs.millis.ai/pricing"}],"sources":["https://docs.millis.ai/pricing","https://frontdeskreview.com/software/ai-voice-agents/millis-ai/"],"confidence":"medium","unverified":"Whether model/TTS list is current; concurrency; free credit; SDK snippet.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":true,"open_source":false,"self_hostable":false,"platform_fee_per_min":0.02,"all_in":false,"phone_numbers":true,"byo_llm":true,"byo_keys":null,"no_code_builder":null,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"$0.02/min platform fee plus STT, LLM and per-character TTS; phone minutes from your own carrier. Pricing page looks stale.","_added":[]}},{"id":"voiceflow","name":"Voiceflow (voice agents)","vendor":"Voiceflow","category":"platform","summary":"Collaborative agent design platform for chat and voice; the official pricing page now shows only 'free trial' and 'request pricing' tiers, with usage-based credits.","status":"GA","models":[{"name":"Agencies & Partners","status":"GA","notes":"Free trial, no credit card; usage-based billing."},{"name":"Businesses","status":"GA","notes":"Request pricing."}],"transports":["PSTN","WebRTC"],"audio":{"input":"Phone via Twilio/Vonage integration, web","output":"Same"},"languages":"Not verified.","latency":"No verified number.","features":["visual flow designer","knowledge base","multi-model","voice + chat","observability"],"pricing":{"model":"usage","items":[{"what":"Plans","price":"Not published on official page","unit":"","notes":"Third-party blogs cite Pro $60/month and Business $150/month with credits, and roughly 170 credits per 10-minute call; unverified."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Official page has no prices; third-party numbers conflict."},"free_tier":"Free trial, no credit card (official).","source":"https://www.voiceflow.com/pricing"},"limits":["Third parties report agents stop responding when credits run out and no mid-cycle top-up (unverified)"],"regions":"Not verified.","setup":{"steps":["Start a free trial, design the agent in the canvas, connect a Twilio number for phone."],"endpoint":"Voiceflow APIs","auth":"API key","snippet_lang":"javascript","snippet":"// Voiceflow agents are designed in the canvas; no minimal voice snippet verified on 2026-10-10."},"warnings":[{"severity":"medium","title":"Pricing opaque","detail":"No public prices on the official page as of 2026-10-10; get a written quote including voice minute consumption."},{"severity":"medium","title":"Credits drain fast on voice","detail":"Voice uses far more credits than chat per third-party analysis; model a month of calls before committing."},{"severity":"low","title":"Telephony separate","detail":"Carrier fees (Twilio/Vonage) are on top."}],"best_for":"Teams already designing chat agents in Voiceflow that want to add a phone channel.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Pricing","url":"https://www.voiceflow.com/pricing"}],"sources":["https://www.voiceflow.com/pricing","https://www.ringly.io/blog/voiceflow-pricing","https://www.getmacha.com/blog/voiceflow-pricing-explained"],"confidence":"low","unverified":"All prices and limits.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"platform_fee_per_min":null,"all_in":null,"phone_numbers":true,"byo_llm":null,"byo_keys":null,"no_code_builder":true,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Official page lists only a free trial and request-pricing tiers; credit-based usage. Third-party plan prices unverified.","_added":[]}},{"id":"thoughtly","name":"Thoughtly","vendor":"Thoughtly","category":"platform","summary":"Sales-oriented AI phone agent platform with flat plans (unlimited voice minutes within concurrency limits) rather than per-minute pricing; all tiers are quote-based on the official page.","status":"GA","models":[{"name":"Flex","status":"GA","notes":"Unlimited voice minutes/SMS/email, 10 concurrent calls, self-serve onboarding, email support. Third parties report it starts at $500/month (not on official page)."},{"name":"Scale","status":"GA","notes":"100+ concurrent calls, 99.9% SLA, auto-QA, TCPA/state DNC compliance review, SOC 2 Type II report on request."},{"name":"Enterprise","status":"GA","notes":"HIPAA + BAA, PCI scope review, data residency, SSO/SCIM; for 1M+ minutes/month or regulated verticals."}],"transports":["PSTN"],"audio":{"input":"Phone","output":"Phone"},"languages":"34+ languages (vendor).","latency":"No verified number.","features":["voice cloning (20-second sample)","SMS/email follow-up","workflows","200+ integrations","two-way CRM sync","REST API","auto-QA (Scale)"],"pricing":{"model":"subscription","items":[{"what":"Flex / Scale / Enterprise","price":"Quote only","unit":"per month","notes":"Usage above plan allowances billed in arrears per FAQ; carrier fees may be extra."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Flat-fee model; effective per-minute cost depends on volume. Third-party example: $500/month is $0.50/min at 1,000 minutes, $0.05/min at 10,000 minutes (unverified base price)."},"free_tier":"Third parties mention a 14-day trial that cannot call external numbers (unverified).","source":"https://www.thoughtly.com/pricing"},"limits":["Flex: 10 concurrent calls","Scale: 100+"],"regions":"Data residency on Enterprise.","setup":{"steps":["Talk to sales or start Flex self-serve onboarding.","Build the agent and workflows in the app, connect CRM, attach numbers."],"endpoint":"Thoughtly REST API","auth":"API key","snippet_lang":"javascript","snippet":"// Thoughtly agents are built in its app; no public minimal API snippet verified on 2026-10-10."},"warnings":[{"severity":"medium","title":"Flat fee punishes low volume","detail":"Unlimited minutes only pay off at high usage; at low volume the effective per-minute price is very high."},{"severity":"medium","title":"HIPAA only on Enterprise","detail":"Flex and Scale do not include a BAA."},{"severity":"low","title":"Overage wording conflicts","detail":"Plan table says unlimited, FAQ says usage above allowances is billed in arrears; get it in writing."}],"best_for":"Sales/inbound teams with steady high call volume wanting predictable monthly cost.","open_source":false,"self_hostable":false,"compliance":"SOC 2 Type II report on request (Scale+); HIPAA + BAA on Enterprise only.","docs":[{"label":"Pricing","url":"https://www.thoughtly.com/pricing"}],"sources":["https://www.thoughtly.com/pricing","https://costbench.com/software/voice-ai-agents/thoughtly/"],"confidence":"medium","unverified":"Dollar prices (not on official page); trial terms.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":34,"max_session_min":null,"concurrency":10,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":true,"soc2":true,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"platform_fee_per_min":null,"all_in":null,"phone_numbers":true,"byo_llm":null,"byo_keys":null,"no_code_builder":null,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Flat quote-based plans with unlimited minutes within concurrency (Flex 10, Scale 100+). Third parties report Flex from $500/month (unverified). Languages is vendor 34+. HIPAA/BAA on Enterprise only; SOC 2 report on Scale+.","_added":[]}},{"id":"playai-agents","name":"PlayAI Agents","vendor":"PlayAI (acquired by Meta)","category":"platform","summary":"Former voice-agent and TTS platform (PlayHT/PlayAI). Meta acquired PlayAI in July 2025 and the team joined Meta; on 2026-10-10 the play.ai domain did not resolve.","status":"Deprecated","models":[{"name":"PlayAI Agents / API","status":"Deprecated (unavailable)","notes":"No reachable product site."}],"transports":[],"audio":{"input":"N/A","output":"N/A"},"languages":"N/A","latency":"N/A","features":[],"pricing":{"model":"usage","items":[],"est_per_minute_usd":{"low":null,"high":null,"basis":"Product not available."},"free_tier":"N/A","source":"https://www.hpcwire.com/aiwire/2025/07/16/meta-picks-up-voice-startup-play-ai-as-it-rebuilds-around-genai/"},"limits":[],"regions":"N/A","setup":{"steps":["Migrate existing PlayAI voices/agents to another provider."],"endpoint":"","auth":"","snippet_lang":"javascript","snippet":"// PlayAI is no longer reachable (play.ai DNS lookup failed on 2026-10-10)."},"warnings":[{"severity":"high","title":"Not available","detail":"play.ai did not resolve on 2026-10-10 after the 2025 Meta acquisition. Platforms that listed PlayHT voices (e.g. Millis docs) may still show them; expect failures."},{"severity":"low","title":"No official sunset notice found","detail":"Status inferred from the acquisition reports and the dead domain."}],"best_for":"Nothing; listed so users know to migrate.","open_source":false,"self_hostable":false,"compliance":"N/A","docs":[],"sources":["https://www.hpcwire.com/aiwire/2025/07/16/meta-picks-up-voice-startup-play-ai-as-it-rebuilds-around-genai/","https://play.ai/"],"confidence":"medium","unverified":"Whether any PlayAI API endpoints still serve legacy customers.","cat":"platforms","kind":"platform","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"platform_fee_per_min":null,"all_in":null,"phone_numbers":null,"byo_llm":null,"byo_keys":null,"no_code_builder":null,"recording":null,"noise_cancellation":null,"turn_detection_model":null,"_notes":"Unavailable: acquired by Meta in 2025; play.ai did not resolve on 2026-10-10.","_added":[]}},{"id":"tavus-cvi","name":"Tavus Conversational Video Interface (CVI)","vendor":"Tavus","category":"avatar","summary":"Hosted realtime video agents ('PALs') with Tavus's own face, perception and turn-taking models. You create a conversation by API and get back a Daily room URL to embed.","status":"GA","models":[{"name":"Phoenix-4.5","status":"GA","notes":"Face rendering model; default for newly trained faces since 9 Sep 2026 per changelog. 50 new Phoenix-4.5 stock faces added 10 Sep 2026."},{"name":"Raven","status":"GA","notes":"Perception model: reads emotion, body language and screen share from the user's camera."},{"name":"Sparrow-2","status":"GA","notes":"Turn-taking model; default turn detection since 8 Sep 2026, set via turn_detection_model."}],"transports":["WebRTC","Daily","LiveKit","Pipecat","Agora"],"audio":{"input":"User microphone in the Daily room (or your own pipeline via LiveKit/Pipecat/Echo modes)","output":"Synthesized voice from Tavus stock voices or integrated TTS"},"languages":"Changelog says stock voices can speak 42 supported languages.","latency":"Vendor describes Sparrow as delivering 'ultra-fast response times'; no number published on the models page.","features":["managed full pipeline (STT, LLM, TTS)","bring your own LLM","custom avatar from video (replica training)","perception of user video and screen share","joins Google Meet, Zoom and Microsoft Teams calls","recordings and closed captions","memory stores per participant","MCP connectors for actions"],"pricing":{"model":"subscription","items":[{"what":"Free","price":"$0","unit":"per month","notes":"20 CVI minutes, 1 concurrent, 5 min max conversation, no overage"},{"what":"Starter","price":"$22","unit":"per month","notes":"60 CVI minutes, 1 concurrent, 10 min max, no pay-as-you-go overage (minutes are a hard cap)"},{"what":"Builder","price":"$59","unit":"per month","notes":"175 minutes, 3 concurrent, 15 min max, overage $0.35/min"},{"what":"Growth","price":"$397","unit":"per month","notes":"1,300 minutes, 10 concurrent, 60 min max, overage $0.31/min"},{"what":"Business","price":"$975","unit":"per month","notes":"4,000 minutes, 15 concurrent, 60 min max, overage $0.26/min"},{"what":"Conversation recordings","price":"$0.03","unit":"per minute","notes":"From an older table on the same page; verify"},{"what":"Billing granularity","price":"30 second minimum","unit":"per conversation","notes":"Then rounded to the nearest 6 seconds"}],"est_per_minute_usd":{"low":0.24,"high":0.37,"basis":"Business $975/4,000 min = $0.24 up to Starter $22/60 min = $0.37; overage $0.26-$0.35"},"free_tier":"Free plan: 20 CVI minutes per month, 1 concurrent session, 5 minute max conversation.","source":"https://www.tavus.io/pricing"},"limits":["Concurrent sessions: 1 (Free/Starter), 3 (Builder), 10 (Growth), 15 (Business)","Max conversation: 5 / 10 / 15 / 60 / 60 minutes by plan; max_call_duration is silently capped to the plan maximum","participant_absent_timeout default 300 s; participant_left_timeout default 0","Custom faces per plan: 1 (Starter), 3 (Builder), 7 (Growth), 15 (Business)"],"regions":"Not published on the pages checked.","setup":{"steps":["Create an API key in the Tavus developer portal.","Pick a stock face_id or train a custom face (needs a consent statement for a real person).","Create or pick a PAL (persona) and note its pal_id.","POST /v2/conversations from your backend with tight timeouts.","Embed the returned conversation_url (a Daily room) or join it with the Daily SDK; end the conversation via API when done."],"endpoint":"POST https://tavusapi.com/v2/conversations","auth":"x-api-key header (server-side only)","snippet_lang":"javascript","snippet":"// Server-side (Node 18+)\nconst res = await fetch(\"https://tavusapi.com/v2/conversations\", {\n  method: \"POST\",\n  headers: {\n    \"x-api-key\": process.env.TAVUS_API_KEY,\n    \"Content-Type\": \"application/json\",\n  },\n  body: JSON.stringify({\n    face_id: \"rc9cff32ceba\",   // stock or custom face\n    pal_id: \"pcb7a34da5fe\",    // your PAL (persona)\n    conversation_name: \"Support demo\",\n    properties: {\n      max_call_duration: 600,         // seconds, capped to plan max\n      participant_left_timeout: 30,   // end soon after user leaves\n      participant_absent_timeout: 60, // end if nobody joins\n    },\n  }),\n});\nconst { conversation_id, conversation_url } = await res.json();\n// Send conversation_url to the browser and open it in an iframe\n// or with the Daily JS SDK. It is a Daily room."},"warnings":[{"severity":"high","title":"Pricing page shows several tables","detail":"tavus.io/pricing contains older tables (e.g. $59 Starter with 100 minutes, $0.37 overage) next to the newer ladder ($22 Starter, $59 Builder). Third-party trackers say the newer ladder is current. Confirm in the dashboard before quoting prices."},{"severity":"high","title":"Starter has no overage, so calls just stop","detail":"Free and $22 Starter have no pay-as-you-go: when included minutes run out, conversations stop. Fine for demos, dangerous for production."},{"severity":"medium","title":"Default participant_left_timeout is 0 but absent timeout is 300 s","detail":"A conversation created but never joined can sit for 5 minutes by default. Set participant_absent_timeout low and always call End Conversation."},{"severity":"medium","title":"API renamed replica/persona to face/PAL","detail":"Current docs use face_id and pal_id; older tutorials use replica_id and persona_id. Check the field names against the live API reference."},{"severity":"medium","title":"Consent needed for personal replicas","detail":"Training a face of a real person needs a spoken consent statement read verbatim; footage of someone else needs their own matching consent video."}],"best_for":"Turnkey customer-facing video agents and meeting bots where you want one vendor for face, perception and turn-taking.","open_source":false,"self_hostable":false,"compliance":"Not verified in this pass; check Tavus trust/security pages for SOC 2, HIPAA and GDPR terms.","docs":[{"label":"Create conversation","url":"https://docs.tavus.io/api-reference/conversations/create-conversation.md"},{"label":"Call duration and timeouts","url":"https://docs.tavus.io/sections/conversational-video-interface/conversation/customizations/call-duration-and-timeout.md"},{"label":"Models","url":"https://docs.tavus.io/sections/models.md"},{"label":"Changelog","url":"https://docs.tavus.io/sections/changelog/changelog.md"}],"sources":["https://www.tavus.io/pricing","https://usagepricing.com/blueprint/tavus","https://docs.tavus.io/sections/changelog/changelog.md","https://docs.tavus.io/sections/replica/replica-training"],"confidence":"medium","unverified":"Which pricing table on tavus.io is live; Free plan custom-face allowance; compliance certifications; regions.","cat":"video","kind":"avatar","verified_at":"2026-10-10","short":"Tavus CVI","facts":{"latency_ms":500,"languages":42,"max_session_min":60,"concurrency":1,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":22,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.24,"custom_avatar":true,"photo_avatar":null,"byo_llm":true,"byo_tts":null,"interruptions":true,"resolution_p":1080,"fps":40,"cpu_ok":null,"_notes":"Latency is vendor claim (under 500 ms average). Max session 5/10/15/60 min by plan. Price is Business $975/4,000 min; overage $0.26-$0.35. Resolution and fps are Phoenix-4 launch figures (current model Phoenix-4.5). Pricing page shows conflicting tables.","_added":["latency_ms: https://tavus.io/cvi","resolution_p: https://chatforest.com/reviews/tavus-conversational-video-ai-api/ (Phoenix-4 launch coverage)","fps: https://chatforest.com/reviews/tavus-conversational-video-ai-api/ (Phoenix-4 launch coverage)"]}},{"id":"heygen-liveavatar","name":"LiveAvatar (formerly HeyGen Interactive / Streaming Avatar)","vendor":"HeyGen","category":"avatar","summary":"HeyGen's realtime avatar platform. FULL mode runs speech recognition, LLM, TTS and WebRTC for you; LITE (Avatar Only) mode just renders the face from audio you send.","status":"GA","models":[{"name":"FULL mode","status":"GA","notes":"Managed ASR + LLM + TTS + WebRTC; default managed LLM gpt-5.4-nano since 17 Sep 2026; custom LLM and TTS supported."},{"name":"LITE / Avatar Only mode","status":"GA","notes":"You bring STT, LLM, TTS and the room (LiveKit, Agora) or an ElevenLabs, Cartesia, OpenAI Realtime or Gemini Realtime agent config."},{"name":"HeyGen Streaming Avatar API v1/v2","status":"Deprecated","notes":"HeyGen docs mark the Streaming API deprecated and point to LiveAvatar; Interactive Avatar sunset stated as 31 March 2026."}],"transports":["WebRTC","LiveKit","WebSocket","Agora"],"audio":{"input":"FULL: user mic via managed WebRTC. LITE: you send PCM 16-bit 24 kHz base64 chunks over WebSocket (about 1 s chunks, max 1 MB per packet)","output":"FULL: TTS voices (ElevenLabs, Cartesia, Fish Audio options). LITE: lip-synced to your audio"},"languages":"A Get Languages endpoint exists; list not captured in this pass.","latency":"No vendor latency figure captured.","features":["bring your own LLM","bring your own TTS","lip sync from any audio (LITE)","push-to-talk mode","session memory across sessions","sandbox mode that does not consume credits","iframe embed","transcripts API"],"pricing":{"model":"credits","items":[{"what":"FULL / Embed mode","price":"2 credits","unit":"per minute","notes":"All plans"},{"what":"LITE / Avatar Only mode","price":"1 credit","unit":"per minute","notes":"All plans"},{"what":"Free","price":"$0","unit":"per month","notes":"10 credits, no overage"},{"what":"Essential","price":"$99","unit":"per month","notes":"1,100 (+10) credits; overage $0.095/credit"},{"what":"Business","price":"$475","unit":"per month","notes":"6,000 (+10) credits; overage $0.09/credit"}],"est_per_minute_usd":{"low":0.08,"high":0.19,"basis":"Essential $99/1,110 credits = ~$0.089/credit, Business ~$0.079/credit: LITE 1 credit/min = $0.08-0.09, FULL 2 credits/min = $0.16-0.19 incl. overage"},"free_tier":"10 credits per month (about 5 FULL minutes or 10 LITE minutes); sandbox mode is free for development.","source":"https://docs.liveavatar.com/docs/faq/credits.md"},"limits":["Sessions cannot start without credits for at least one minute","Without overage enabled, sessions end automatically when credits run out","max_session_duration must not exceed your tier's limit (limit values not published in docs checked)","Rate limits: 10 req/s per API key; 5 req/s per session token on start/stop/keep-alive (HTTP 429)"],"regions":"Not published in docs checked; a firewall configuration FAQ exists.","setup":{"steps":["Sign up at app.liveavatar.com and copy an API key from the Developers page.","Pick a public avatar or one migrated from HeyGen.","Backend: POST /v1/sessions/token with mode FULL or LITE (use is_sandbox: true while developing).","Backend: POST /v1/sessions/start with that token and pass the returned WebRTC credentials to the browser.","Frontend: join with the LiveAvatar Web SDK (or your LiveKit room in LITE mode); stop the session via API when done."],"endpoint":"https://api.liveavatar.com/v1/sessions/token","auth":"X-API-KEY header for minting session tokens; session token for start/stop","snippet_lang":"javascript","snippet":"// Server-side: mint a session token, then start the session\nconst H = {\n  \"X-API-KEY\": process.env.LIVEAVATAR_API_KEY,\n  \"Content-Type\": \"application/json\",\n};\nconst tokRes = await fetch(\"https://api.liveavatar.com/v1/sessions/token\", {\n  method: \"POST\",\n  headers: H,\n  body: JSON.stringify({\n    mode: \"FULL\",              // or \"LITE\" for avatar-only\n    avatar_id: process.env.AVATAR_ID,\n    max_session_duration: 600, // seconds\n    interactivity_type: \"CONVERSATIONAL\",\n    is_sandbox: true,          // no credits while developing\n  }),\n});\nconst { data } = await tokRes.json();\n// data.session_id, data.session_token\n// Next: POST https://api.liveavatar.com/v1/sessions/start with the\n// session token (see Start Session reference for the exact auth\n// form), then hand the WebRTC credentials to the Web SDK.\n// github.com/heygen-com/liveavatar-web-sdk"},"warnings":[{"severity":"high","title":"Old HeyGen streaming code is dead","detail":"HeyGen's Streaming/Interactive Avatar API is deprecated with a stated sunset of 31 March 2026, and avatar versions v1/v2 return errors. Migrate to LiveAvatar; avatars and knowledge bases are copied automatically but individual voices are not migrated."},{"severity":"high","title":"Overage on means no automatic stop","detail":"With overage enabled, sessions do not end when credits run out and you keep being billed per credit. Always pass max_session_duration and call Stop Session."},{"severity":"medium","title":"FULL mode costs double","detail":"FULL is 2 credits/min versus 1 for LITE. If you already run a voice agent (LiveKit, ElevenLabs, OpenAI Realtime), LITE halves the avatar bill."},{"severity":"medium","title":"VP8 is deprecated","detail":"Since 22 Jul 2026 VP8 encoding is deprecated in favour of H264 and will be removed; set video_settings.encoding to H264."},{"severity":"low","title":"Third-party price lists disagree","detail":"Some comparison sites quote a $19/150 credit plan or $0.05/sec API wallet pricing. The LiveAvatar docs credits FAQ (Free / $99 / $475) is the source used here."}],"best_for":"Teams that want HeyGen-quality stock avatars with a cheap avatar-only mode on top of an existing voice agent.","open_source":false,"self_hostable":false,"compliance":"Not verified in this pass.","docs":[{"label":"Overview","url":"https://docs.liveavatar.com/"},{"label":"Credits and subscriptions","url":"https://docs.liveavatar.com/docs/faq/credits.md"},{"label":"Create session token","url":"https://docs.liveavatar.com/api-reference/sessions/create-session-token.md"},{"label":"LITE mode lifecycle","url":"https://docs.liveavatar.com/docs/lite-mode/lifecycle.md"},{"label":"Changelog","url":"https://docs.liveavatar.com/changelog.md"},{"label":"HeyGen streaming deprecation","url":"https://docs.heygen.com/docs/streaming-api-deprecated"}],"sources":["https://docs.liveavatar.com/docs/faq/credits.md","https://docs.liveavatar.com/api-reference/sessions/create-session-token.md","https://docs.liveavatar.com/changelog.md","https://docs.liveavatar.com/docs/interactive-avatar-migration-guide"],"confidence":"high","unverified":"Concurrency limits per plan, per-tier max session duration, exact auth for /v1/sessions/start, latency.","cat":"video","kind":"avatar","verified_at":"2026-10-10","short":"LiveAvatar (HeyGen)","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":99,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.08,"custom_avatar":null,"photo_avatar":null,"byo_llm":true,"byo_tts":true,"interruptions":null,"resolution_p":720,"fps":null,"cpu_ok":null,"_notes":"Price is LITE mode (1 credit/min) at Business credit rate; FULL mode is 2 credits/min (about $0.16-$0.19). Free plan 10 credits/month. Old HeyGen Streaming API deprecated.","_added":["resolution_p: https://docs.agora.io/en/ai/models/avatar/heygen (quality high = 720p)"]}},{"id":"d-id-agents","name":"D-ID Agents (realtime streams)","vendor":"D-ID","category":"avatar","summary":"D-ID's realtime talking-head agents, now built on its V4 Expressive model. Agents Streams V2 runs on LiveKit; there is a client SDK and an embed script.","status":"GA","models":[{"name":"V4 Expressive Visual Agents","status":"GA","notes":"Launched 16 Mar 2026; diffusion-based, sentiment-aware expressions. Vendor claims sub-0.5 s conversational turns and up to 4K output."},{"name":"Agents Streams V2 (LiveKit)","status":"GA","notes":"Only for Expressive Agents; commands and events over the LiveKit data channel."},{"name":"Agents Streams V1 (WebRTC)","status":"GA","notes":"Used by other agent types; SDP/ICE exchange via REST."}],"transports":["WebRTC","LiveKit"],"audio":{"input":"User speech or text chat via SDK","output":"D-ID managed TTS / voice clones"},"languages":"Not verified in this pass.","latency":"Vendor claim: sub-0.5 s conversational turns (V4). One developer forum report measured about 2.1 s extra before video starts in the browser.","features":["managed LLM agent with knowledge base","LiveKit SDK path for backend orchestration","embed script with no backend","expressive sentiments","custom avatars"],"pricing":{"model":"credits","items":[{"what":"Trial","price":"$0","unit":"trial","notes":"Up to 10 streaming minutes (third-party reproduction of API table)"},{"what":"Build","price":"$14.40","unit":"per month (annual)","notes":"64 credits, up to 32 streaming minutes, personal use only"},{"what":"Launch","price":"$35","unit":"per month","notes":"180 credits, up to 90 streaming minutes; commercial use starts here"},{"what":"Scale","price":"$138.60","unit":"per month","notes":"800 credits, up to 400 streaming minutes"},{"what":"Enterprise","price":"Custom","unit":"","notes":"Concurrency and fastest processing appear to be enterprise features"}],"est_per_minute_usd":{"low":0.35,"high":0.45,"basis":"Scale $138.60/400 streaming min = $0.35 to Build $14.40/32 = $0.45; streaming shares the same credit pool as offline video. Third-party data, unverified."},"free_tier":"Trial allowance (reported up to 10 streaming minutes).","source":"https://creatify.ai/blog/d-id-pricing-(2026)-plans-credits-and-what-you-ll-actually-pay"},"limits":["Concurrent stream limit per plan not published (enterprise feature per third-party guide)","Credits do not roll over (third-party)","Client keys are restricted to allowed_domains"],"regions":"Not verified.","setup":{"steps":["Create an agent in D-ID Studio and copy its agent ID.","Backend: POST /agents/client-key with your allowed_domains to get a client key.","Frontend: npm install @d-id/client-sdk and create an agent manager with the client key.","Call connect(), then chat()/speak(), and disconnect() when the user leaves."],"endpoint":"POST https://api.d-id.com/agents/client-key","auth":"Authorization: Basic <API key> on the server; client key in the browser","snippet_lang":"javascript","snippet":"// Browser, after your server created a client key via\n// POST https://api.d-id.com/agents/client-key\nimport * as did from \"@d-id/client-sdk\";\n\nconst video = document.getElementById(\"agent-video\"); // autoplay playsinline\nconst agent = await did.createAgentManager(\"agt_abc123\", {\n  auth: { type: \"key\", clientKey: CLIENT_KEY },\n  callbacks: {\n    onSrcObjectReady(stream) { video.srcObject = stream; },\n    onConnectionStateChange(state) { console.log(\"state\", state); },\n    onNewMessage(messages) { console.log(messages); },\n  },\n});\n\nawait agent.connect();\nawait agent.chat(\"Hello! What can you help me with?\");\n// ...\nawait agent.disconnect(); // stop billing"},"warnings":[{"severity":"high","title":"Official API pricing page did not render","detail":"d-id.com/pricing/api returned only navigation to our fetcher; plan numbers here come from a third-party reproduction. Confirm in the D-ID dashboard."},{"severity":"high","title":"Cheapest plans are personal-use only","detail":"Per the third-party API table, Trial and Build carry a personal-use licence; commercial API use starts at Launch."},{"severity":"medium","title":"Streaming eats the same credits as video","detail":"Streaming minutes are drawn from the same credit pool as offline renders, so the 'up to N streaming minutes' figures are not additive."},{"severity":"medium","title":"Measure latency yourself","detail":"Vendor claims sub-0.5 s turns, but a developer reported about 2.1 s before playback starts. Test end-to-end in your own stack."},{"severity":"low","title":"Two stream APIs","detail":"V2 (LiveKit) only works with Expressive Agents; other agents still use V1 WebRTC endpoints. Do not mix tutorials."}],"best_for":"Marketing and training sites that want an embeddable expressive agent with minimal backend work.","open_source":false,"self_hostable":false,"compliance":"Not verified in this pass.","docs":[{"label":"Agent sessions quickstart","url":"https://docs.d-id.com/docs/agent-session-quickstart.md"},{"label":"LiveKit overview","url":"https://docs.d-id.com/docs/livekit-overview.md"},{"label":"Docs index","url":"https://docs.d-id.com/llms.txt"},{"label":"API pricing","url":"https://www.d-id.com/pricing/api/"}],"sources":["https://docs.d-id.com/docs/agent-session-quickstart.md","https://docs.d-id.com/docs/livekit-overview.md","https://trainingindustry.com/press-release/artificial-intelligence/d-id-launches-v4-expressive-visual-agents-for-real-time-llm-connected-interaction-at-enterprise-scale/","https://creatify.ai/blog/d-id-pricing-(2026)-plans-credits-and-what-you-ll-actually-pay","https://docs.d-id.com/discuss/6850034dbb468f006953f084"],"confidence":"low","unverified":"All plan prices and streaming-minute allowances (third-party), concurrency, languages, session limits.","cat":"video","kind":"avatar","verified_at":"2026-10-10","short":"D-ID Agents","facts":{"latency_ms":500,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":14.4,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.35,"custom_avatar":true,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"Latency is vendor claim (sub-0.5 s turns); a developer measured about 2.1 s before video starts. Plan prices are third-party: Build $14.40/month (annual, personal use only), commercial from Launch $35. Price is Scale $138.60/400 min.","_added":[]}},{"id":"simli","name":"Simli","vendor":"Simli","category":"avatar","summary":"Low-latency audio-to-video face API: you stream your agent's audio in and get a lip-synced face back over WebRTC or LiveKit. Also offers a managed 'Simli Auto' agent.","status":"GA","models":[{"name":"Trinity","status":"GA","notes":"Current face model with its own face generation endpoints; described by third parties as newer and cheaper than Legacy."},{"name":"Legacy faces","status":"GA","notes":"Older face model, still has endpoints."},{"name":"Gaussian / new facial model","status":"Preview","notes":"Homepage mentions a new high-resolution facial model and Gaussian models demo."}],"transports":["WebRTC","LiveKit","WebSocket"],"audio":{"input":"PCM16 audio from your TTS (audioInputFormat pcm16)","output":"None added; your audio is passed through with the video"},"languages":"Language-agnostic lip sync from audio (not explicitly verified).","latency":"Vendor claim: under 300 ms for the speech-to-video stage.","features":["lip sync from any audio","bring your own LLM and TTS","custom face from a photo","LiveKit and Pipecat plugins","handleSilence idle animation","active session count endpoint"],"pricing":{"model":"per-minute","items":[{"what":"Free plan","price":"$10 credit + 50 min/month","unit":"signup credit plus monthly top-up","notes":"From simli.com homepage"},{"what":"Pay-as-you-go","price":"about $0.05","unit":"per minute","notes":"Third-party directories only; not on an official page we could load"}],"est_per_minute_usd":{"low":0.05,"high":0.05,"basis":"Third-party reported PAYG rate; unverified"},"free_tier":"$10 on signup and a monthly top-up of 50 minutes (homepage).","source":"https://www.simli.com/"},"limits":["maxSessionLength default 3600 s","maxIdleTime default 300 s","Concurrency per plan not published (Active Session Count endpoint exists)"],"regions":"Not published.","setup":{"steps":["Get an API key from the Simli dashboard.","Pick a preset face or generate a Trinity face from an image.","Backend: POST /compose/token to get a session token (set short maxSessionLength and maxIdleTime).","Frontend or agent: connect via WebRTC SDK or the LiveKit plugin and stream PCM16 audio from your TTS."],"endpoint":"POST https://api.simli.ai/compose/token","auth":"x-simli-api-key header","snippet_lang":"javascript","snippet":"// Server-side: get a short-lived session token\nconst res = await fetch(\"https://api.simli.ai/compose/token\", {\n  method: \"POST\",\n  headers: {\n    \"x-simli-api-key\": process.env.SIMLI_API_KEY,\n    \"Content-Type\": \"application/json\",\n  },\n  body: JSON.stringify({\n    faceId: process.env.SIMLI_FACE_ID,\n    handleSilence: true,\n    maxSessionLength: 600, // default 3600\n    maxIdleTime: 60,       // default 300\n    audioInputFormat: \"pcm16\",\n  }),\n});\nconst { session_token } = await res.json();\n// Give session_token to the Simli client SDK (WebRTC) or use the\n// LiveKit plugin, then stream PCM16 audio from your TTS into it."},"warnings":[{"severity":"high","title":"No official price list found","detail":"simli.com shows the free tier but no per-minute rate; the ~$0.05/min figure is from third-party directories. Confirm in the dashboard before committing."},{"severity":"medium","title":"Defaults allow 1 hour sessions and 5 minute idle","detail":"maxSessionLength defaults to 3600 s and maxIdleTime to 300 s; lower both so abandoned tabs stop billing."},{"severity":"medium","title":"You own the rest of the pipeline","detail":"Simli only renders the face; STT, LLM, TTS and the room are separate bills and separate latency."},{"severity":"low","title":"Two face model families","detail":"Trinity and Legacy faces use different creation endpoints; make sure your faceId matches the model you intend."}],"best_for":"Developers who already have a voice agent and want the cheapest, fastest face layer.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Docs index","url":"https://docs.simli.com/llms.txt"},{"label":"Compose session token","url":"https://docs.simli.com/api-reference/compose-session-token.md"},{"label":"LiveKit","url":"https://docs.simli.com/api-reference/livekit.md"}],"sources":["https://www.simli.com/","https://docs.simli.com/api-reference/compose-session-token.md","https://www.voiceaispace.com/tool/simli"],"confidence":"low","unverified":"Per-minute price, paid plans, concurrency, Trinity-specific pricing.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":300,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":10,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.05,"custom_avatar":true,"photo_avatar":true,"byo_llm":true,"byo_tts":true,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"Latency is the vendor claim for the speech-to-video stage only. Price about $0.05/min is from third-party directories. Free: $10 signup plus 50 min/month. maxSessionLength default 60 min (configurable).","_added":[]}},{"id":"anam","name":"Anam","vendor":"Anam","category":"avatar","summary":"Realtime photoreal personas with a managed voice + LLM pipeline or your own LLM. Clear published per-minute pricing with per-second billing.","status":"GA","models":[{"name":"cara-4","status":"GA","notes":"Recommended for new production avatars; 1152x768 landscape or 768x1152 portrait."},{"name":"cara-3","status":"GA","notes":"Previous stable model kept for existing deployments; 720x480."},{"name":"cara-4-latest","status":"Preview","notes":"Experimental; vendor says not for production sessions."}],"transports":["WebRTC","LiveKit"],"audio":{"input":"User mic via Anam JS/Python SDK","output":"Anam voices including custom voice, or audio passthrough"},"languages":"Multilingual support page exists; list not captured.","latency":"Not captured; Anam publishes a session performance page.","features":["managed full pipeline","bring your own LLM (CUSTOMER_CLIENT_V1 client-side LLM)","custom avatar from photo in under 2 minutes (vendor)","custom voice on all plans","LiveKit integration","spend cap setting","concurrency endpoint"],"pricing":{"model":"subscription","items":[{"what":"Free","price":"$0","unit":"per month","notes":"30 min, 1 concurrent, 1 custom avatar, 3 min conversation limit, no overage"},{"what":"Starter","price":"$12","unit":"per month","notes":"50 min, overage $0.16/min, 1 concurrent, 5 min limit"},{"what":"Explorer","price":"$49","unit":"per month","notes":"250 min, overage $0.14/min, 3 concurrent, 10 min limit"},{"what":"Growth","price":"$299","unit":"per month","notes":"2,000 min, overage $0.12/min, 5 concurrent, 2 h limit"},{"what":"Professional","price":"$999","unit":"per month","notes":"8,000 min, overage $0.11/min, 10 concurrent, 2 h limit"},{"what":"Enterprise","price":"$0.04","unit":"per minute","notes":"Contact sales; 100 to unlimited concurrency"}],"est_per_minute_usd":{"low":0.11,"high":0.24,"basis":"Overage $0.11-$0.16; Starter $12/50 min = $0.24; Enterprise listed at $0.04"},"free_tier":"30 minutes per month, 1 concurrent session, 3 minute conversations.","source":"https://anam.ai/pricing"},"limits":["Concurrent: 1 / 1 / 3 / 5 / 10 by plan; shared across all your sites","Conversation length: 3 min (Free), 5 min (Starter), 10 min (Explorer), 2 h (Growth/Professional)","Session token valid 1 hour","Unused minutes expire monthly"],"regions":"Not captured.","setup":{"steps":["Create an API key in Anam Lab.","Pick avatarId, voiceId and llmId (or a saved personaId).","Backend: POST /v1/auth/session-token with the persona config.","Frontend: create the Anam client with the session token and stream to a video element.","Stop streaming when the user leaves; set a spend cap in Account Settings."],"endpoint":"POST https://api.anam.ai/v1/auth/session-token","auth":"Authorization: Bearer <API key> on server; session token (JWT, 1 h) in browser","snippet_lang":"javascript","snippet":"// Server-side\nconst r = await fetch(\"https://api.anam.ai/v1/auth/session-token\", {\n  method: \"POST\",\n  headers: {\n    Authorization: `Bearer ${process.env.ANAM_API_KEY}`,\n    \"Content-Type\": \"application/json\",\n  },\n  body: JSON.stringify({\n    personaConfig: {\n      name: \"Cara\",\n      avatarId: \"071b0286-4cce-4808-bee2-e642f1062de3\",\n      voiceId: \"de23e340-1416-4dd8-977d-065a7ca11697\",\n      llmId: \"a7cf662c-2ace-4de1-a21e-ef0fbf144bb7\",\n      systemPrompt: \"You are a helpful assistant.\",\n      maxSessionLengthSeconds: 600,\n    },\n  }),\n});\nconst { sessionToken } = await r.json();\n\n// Browser (npm i @anam-ai/js-sdk)\nimport { createClient } from \"@anam-ai/js-sdk\";\nconst client = createClient(sessionToken);\nawait client.streamToVideoElement(\"persona-video\");\n// later: await client.stopStreaming();"},"warnings":[{"severity":"high","title":"Idle time is billed","detail":"Anam states silence, a muted mic or an idle avatar does not pause an active session; billing is per second until the session ends."},{"severity":"medium","title":"Short caps on low plans","detail":"Free conversations end at 3 minutes and Starter at 5; Explorer at 10. Plan for reconnects or upgrade for longer calls."},{"severity":"medium","title":"Avatar-only config is rejected","detail":"A session token with no voiceId and no audio passthrough on the default WebRTC transport returns HTTP 400."},{"severity":"medium","title":"429 can mean several things","detail":"A rejected start may be your org concurrency limit, vendor capacity or your spend cap. Read the error reason and back off instead of retry-looping."},{"severity":"low","title":"Do not ship cara-4-latest","detail":"It is experimental and not recommended for production; pin cara-4."}],"best_for":"Product teams that want predictable per-minute pricing and a polished managed persona with optional custom LLM.","open_source":false,"self_hostable":false,"compliance":"Not verified in this pass.","docs":[{"label":"Docs index","url":"https://anam.ai/docs/llms.txt"},{"label":"Create session token","url":"https://anam.ai/docs/api-reference/sessions/create-session-token.md"},{"label":"Models","url":"https://anam.ai/docs/introduction/models.md"},{"label":"Billing and limits","url":"https://anam.ai/docs/resources/billing-and-limits.md"}],"sources":["https://anam.ai/pricing","https://anam.ai/docs/introduction/models.md","https://anam.ai/docs/resources/billing-and-limits.md","https://anam.ai/docs/api-reference/sessions/create-session-token.md"],"confidence":"high","unverified":"JS SDK method names (createClient / streamToVideoElement) taken from SDK conventions, not re-read this pass; latency figures; regions.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":120,"concurrency":1,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":12,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.11,"custom_avatar":true,"photo_avatar":true,"byo_llm":true,"byo_tts":true,"interruptions":null,"resolution_p":768,"fps":null,"cpu_ok":null,"_notes":"Conversation cap 3/5/10/120/120 min by plan. Price is Professional overage $0.11 (Enterprise listed $0.04). cara-4 is 1152x768 landscape or 768x1152 portrait. Idle time is billed.","_added":[]}},{"id":"bithuman","name":"bitHuman","vendor":"bitHuman","category":"avatar","summary":"Avatar runtime that can render on your own CPU-only machines (no GPU) as well as in bitHuman's cloud, billed by credits per active minute.","status":"GA","models":[{"name":"Essence 2","status":"GA","notes":"A real person from one portrait, up to 1080p; 500 credits to create."},{"name":"Expression 2","status":"GA","notes":"Any character (cartoon, animal, robot, human) from one photo, motion generated from audio; 2,000 credits to create, about 2-2.5 hours."},{"name":"Essence 2 Max","status":"Deprecated","notes":"Retired; use Essence 2."}],"transports":["LiveKit","WebRTC"],"audio":{"input":"Audio from your agent/TTS (or bitHuman managed voice chat)","output":"Your TTS or managed voice"},"languages":"Language-agnostic lip sync from audio (not explicitly verified).","latency":"No latency number published; CPU benchmark shows 1.2x to 2.4x real-time rendering on an Intel i7-13700F.","features":["self-hosted CPU rendering","runs on iPhone, Mac, Android, browser, Linux","survives network drops up to 5 minutes","lip sync from any audio","custom avatar from one photo","LiveKit plugin"],"pricing":{"model":"credits","items":[{"what":"Self-hosted (your hardware)","price":"2 credits ($0.02)","unit":"per minute","notes":"Creator plan or higher"},{"what":"bitHuman cloud avatar","price":"4 credits ($0.04)","unit":"per minute","notes":""},{"what":"Managed voice chat, all-inclusive","price":"10 credits ($0.10)","unit":"per minute","notes":""},{"what":"Top-up","price":"$1","unit":"per 100 credits","notes":"Creator+; top-up credits never expire"},{"what":"Creator / Pro / Business / Enterprise","price":"$20 / $99 / $299 / $999","unit":"per month","notes":"2,000 / 10,000 / 50,000 / 250,000 credits; 3 / 10 / 50 / 200 concurrent cloud sessions"}],"est_per_minute_usd":{"low":0.02,"high":0.1,"basis":"Credit rates at the $1 = 100 credits top-up price: self-hosted $0.02, cloud $0.04, all-inclusive voice $0.10"},"free_tier":"Free plan: 99 credits, 1 concurrent cloud session, chat with featured avatars; page says API access on Free only 'until 12 October'.","source":"https://www.bithuman.ai/pricing"},"limits":["Concurrent cloud sessions: 1 / 3 / 10 / 50 / 200 by plan","Self-hosted sessions not concurrency-limited on Creator+ (per page)","A session needs network at start; offline only for up to 5 minutes"],"regions":"Cloud regions not published; self-hosting runs anywhere.","setup":{"steps":["Subscribe to Creator or higher and get an API secret.","Create or pick an avatar and download its .imx model file.","pip install \"bithuman[expression-2]\" and set BITHUMAN_API_SECRET.","Render frames from audio locally, or use the LiveKit plugin to publish the avatar into a room."],"endpoint":"https://api.bithuman.ai/v1/agent/{agent_id}/model/download","auth":"BITHUMAN_API_SECRET environment variable","snippet_lang":"python","snippet":"# pip install \"bithuman[expression-2]\"\n# export BITHUMAN_API_SECRET=\"<your API secret>\"\n# curl -fL -o wise-pup.imx \\\n#   \"https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2\"\nimport bithuman\n\n# Render a lip-synced avatar locally on CPU from an audio file\nwith bithuman.open(\"wise-pup.imx\") as avatar:\n    frames = 0\n    for frame in avatar.render(\"speech.wav\"):\n        frames += 1  # push each frame to your video sink\n    print(frames, \"frames\")\n\n# For live agents use the LiveKit plugin (livekit-agents[bithuman])\n# and start the avatar session before your AgentSession."},"warnings":[{"severity":"high","title":"Free API access has an end date","detail":"The pricing page lists Free-plan API access only 'until 12 October' (2026 presumably). Plan on Creator ($20/mo) for any API work."},{"severity":"high","title":"Self-hosted is not offline-licensed","detail":"Even on your own CPU, sessions check credentials at start and report usage online; rendering only survives network drops of up to 5 minutes."},{"severity":"medium","title":"Idle minutes are billed","detail":"Billing counts active session time whether talking or idle, to the second."},{"severity":"medium","title":"Avatar creation costs credits and time","detail":"Expression 2 costs 2,000 credits and takes about 2-2.5 hours; Essence 2 costs 500 credits."},{"severity":"low","title":"Intel Macs unsupported for CPU mode","detail":"CPU deployment supports x86_64/arm64 Linux and Windows 11 with Python; Intel Macs are not supported."}],"best_for":"Kiosks, on-device and cost-sensitive deployments where you want to avoid GPU bills.","open_source":false,"self_hostable":true,"compliance":"Vendor states usage reports contain no audio, video, images or conversation text.","docs":[{"label":"Pricing","url":"https://docs.bithuman.ai/pricing.md"},{"label":"CPU deployment","url":"https://docs.bithuman.ai/deploy/cpu.md"},{"label":"Models","url":"https://docs.bithuman.ai/models.md"},{"label":"LiveKit","url":"https://docs.bithuman.ai/platforms/livekit.md"}],"sources":["https://www.bithuman.ai/pricing","https://docs.bithuman.ai/deploy/cpu.md","https://docs.bithuman.ai/models.md"],"confidence":"high","unverified":"Latency, the year on the Free API cut-off, yearly prices.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":3,"free_tier":true,"free_credit_usd":0.99,"entry_plan_usd_month":20,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":true,"price_per_min":0.02,"custom_avatar":true,"photo_avatar":true,"byo_llm":true,"byo_tts":true,"interruptions":null,"resolution_p":1080,"fps":null,"cpu_ok":true,"_notes":"$0.02/min self-hosted on your CPU, $0.04 cloud, $0.10 all-inclusive voice. Free plan 99 credits; Free API access listed only until 12 October. Concurrency is cloud sessions on Creator. Self-hosted still needs online license check.","_added":[]}},{"id":"beyond-presence","name":"Beyond Presence","vendor":"Beyond Presence","category":"avatar","summary":"Hyper-realistic avatars sold two ways: speech-to-video (face only, plug into your LiveKit agent) and managed conversational video agents. Prices in EUR.","status":"GA","models":[{"name":"Speech-to-Video","status":"GA","notes":"Avatar layer for your own voice agent; 50 credits per minute."},{"name":"Conversational Video Agents (managed)","status":"GA","notes":"Full managed agent; 100 credits per minute."}],"transports":["LiveKit","WebRTC"],"audio":{"input":"Your agent's audio (speech-to-video) or user mic (managed agent)","output":"Your TTS or managed voice"},"languages":"Language support page exists; list not captured.","latency":"Not captured.","features":["LiveKit plugin (Python and Node)","managed agent option","custom avatars","1080p stock avatars"],"pricing":{"model":"subscription","items":[{"what":"Free","price":"$0","unit":"per month","notes":"40 speech-to-video min or 20 agent min, 1 concurrent, 3 min sessions"},{"what":"Starter","price":"$49 (EUR 49)","unit":"per month","notes":"280 S2V min / 140 agent min; overage EUR 0.175 / EUR 0.35 per min; 10 concurrent; 1 custom avatar"},{"what":"Growth","price":"$149 (EUR 149)","unit":"per month","notes":"1,490 / 745 min; overage EUR 0.10 / EUR 0.20; 25 concurrent; 3 custom avatars"},{"what":"Scale","price":"$349 (EUR 349)","unit":"per month","notes":"4,000 / 2,000 min; overage EUR 0.0875 / EUR 0.175; 50 concurrent; 10 custom avatars"},{"what":"Enterprise","price":"Custom","unit":"","notes":"Vendor says rates fall to EUR 0.03/min at scale"}],"est_per_minute_usd":{"low":0.09,"high":0.35,"basis":"USD plan price / included minutes: Scale $349/4,000 S2V min = $0.087 up to Starter $49/140 agent min = $0.35. Overage is quoted in EUR only."},"free_tier":"40 speech-to-video minutes (or 20 agent minutes) per month, 1 concurrent, 3 minute session cap.","source":"https://www.beyondpresence.ai/pricing"},"limits":["Concurrent sessions: 1 / 10 / 25 / 50 (Free/Starter/Growth/Scale); over the limit returns HTTP 429","Out of minutes without usage-based billing returns HTTP 402","Free sessions capped at 3 minutes; paid agent sessions unlimited length"],"regions":"EU-based vendor; hosting regions not captured.","setup":{"steps":["Get an API key and copy an avatar ID from Beyond Presence Studio.","Start from the LiveKit Agents Python starter.","pip install livekit-plugins-bey and set BEY_API_KEY, BEY_AVATAR_ID and LiveKit credentials.","Start the avatar session before session.start()."],"endpoint":"LiveKit plugin (livekit-plugins-bey); REST: Create Speech-to-Video Session","auth":"BEY_API_KEY","snippet_lang":"python","snippet":"# pip install livekit-plugins-bey\n# env: BEY_API_KEY, BEY_AVATAR_ID, LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET\nimport os\nfrom livekit.plugins import bey\n\n# inside your LiveKit agent entrypoint, after building `session`\n# session = AgentSession(stt=..., llm=..., tts=...)\n\navatar = bey.AvatarSession(avatar_id=os.environ[\"BEY_AVATAR_ID\"])\nawait avatar.start(session, room=ctx.room)\n\n# then start your agent as usual\n# await session.start(agent=MyAgent(), room=ctx.room)"},"warnings":[{"severity":"medium","title":"Overage is in euros","detail":"Plans show USD monthly prices but per-minute overage only in EUR; your card will be charged with FX exposure."},{"severity":"medium","title":"Managed agent minutes cost double","detail":"Conversational agents use 100 credits/min versus 50 for speech-to-video, so the same plan gives half the minutes."},{"severity":"medium","title":"Usage-based billing is opt-in","detail":"Without it, sessions fail with 402 once included minutes are used; with it, there is no automatic stop. Monitor usage."},{"severity":"low","title":"Idle billing undocumented","detail":"Docs do not say whether idle sessions count; assume wall-clock billing and close sessions promptly."}],"best_for":"EU teams on LiveKit that want generous concurrency (10+) on an entry paid plan.","open_source":false,"self_hostable":false,"compliance":"Not verified in this pass.","docs":[{"label":"Speech-to-video quickstart","url":"https://docs.bey.dev/get-started/quickstart/speech-to-video.md"},{"label":"Concurrency and quotas","url":"https://docs.bey.dev/production/concurrency.md"},{"label":"Docs index","url":"https://docs.bey.dev/llms.txt"}],"sources":["https://www.beyondpresence.ai/pricing","https://docs.bey.dev/production/concurrency.md","https://docs.bey.dev/get-started/quickstart/speech-to-video.md"],"confidence":"high","unverified":"Latency, model names, regions, compliance.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":10,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":49,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.087,"custom_avatar":true,"photo_avatar":null,"byo_llm":true,"byo_tts":true,"interruptions":null,"resolution_p":1080,"fps":null,"cpu_ok":null,"_notes":"Price is Scale $349/4,000 speech-to-video min; managed agent minutes cost double. Overage billed in EUR. Free sessions capped at 3 min, paid agent sessions unlimited. 1080p refers to stock avatars.","_added":[]}},{"id":"lemonslice","name":"LemonSlice","vendor":"LemonSlice","category":"avatar","summary":"Realtime video avatars from a single image, including cartoons and non-human characters, with a prompt that shapes movement and expression.","status":"GA","models":[{"name":"LemonSlice-2.1","status":"GA","notes":"Self-serve plans."},{"name":"LemonSlice-2.1 Flash, CWM-1, CWM-1 Lite, Cartoon","status":"GA","notes":"Ultra plan and above."},{"name":"LemonSlice 2.1 Pro","status":"GA","notes":"Enterprise only, high-resolution installations."}],"transports":["LiveKit","WebRTC"],"audio":{"input":"Your agent's TTS audio (self-managed) or hosted speech models","output":"Your TTS or hosted voice"},"languages":"Not captured.","latency":"Not captured.","features":["avatar from any single image","agent_prompt to steer movement and expression","idle prompt and idle timeout","hosted widget","LiveKit plugin (Python and Node)"],"pricing":{"model":"subscription","items":[{"what":"Starter","price":"$8","unit":"per month","notes":"41 free minutes, extra $0.22/min, 3 concurrent, 30 min calls"},{"what":"Creator / Professional / Scale","price":"$40 / $100 / $240","unit":"per month","notes":"Annual-equivalent $33 / $83 / $200"},{"what":"Ultra","price":"$960","unit":"per month","notes":"Annual-equivalent $800; up to 100 concurrent; 2 h calls; extra models"},{"what":"Hosted avatars (widget or hosted API)","price":"+$0.09","unit":"per minute","notes":"Includes speech models"},{"what":"Headline rate","price":"as low as $0.048","unit":"per minute","notes":"At volume"}],"est_per_minute_usd":{"low":0.048,"high":0.31,"basis":"Vendor floor $0.048/min up to Starter overage $0.22 + $0.09 hosted"},"free_tier":"Free plan only lets you chat with featured avatars; API access needs a paid plan.","source":"https://lemonslice.com/pricing"},"limits":["Call length: 30 min (self-serve), 2 h (Ultra), 24 h (Enterprise)","Concurrency: 3 (Starter), up to 100 (Ultra), 1,000+ (Enterprise)","Calls above the concurrency limit billed at a 2x burst multiplier","Output 368x560 via LiveKit plugin, images center-cropped"],"regions":"Not captured.","setup":{"steps":["Subscribe to a paid plan and get an API key.","Install the LiveKit plugin and set LEMONSLICE_API_KEY.","Pass a public image URL (or agent_id) and an optional agent_prompt.","Start the avatar before your AgentSession."],"endpoint":"LiveKit plugin (livekit-agents[lemonslice])","auth":"LEMONSLICE_API_KEY","snippet_lang":"python","snippet":"# uv add \"livekit-agents[lemonslice]~=1.8\"\n# env: LEMONSLICE_API_KEY + LiveKit credentials\nfrom livekit.plugins import lemonslice\n\n# inside your agent entrypoint, after creating `session`\navatar = lemonslice.AvatarSession(\n    agent_image_url=\"https://example.com/face.png\",  # public image/*\n    agent_prompt=\"Be expressive in your movements and use your hands while talking.\",\n    idle_timeout=60,  # seconds; negative disables\n)\nawait avatar.start(session, room=ctx.room)\n# await session.start(...)"},"warnings":[{"severity":"high","title":"Bursting over the limit costs double","detail":"Calls above your plan's concurrency are billed at a 2x multiplier rather than rejected; watch spikes."},{"severity":"medium","title":"Hosted mode adds $0.09/min","detail":"Using the hosted widget or hosted API (with speech models) adds $0.09/min on top of avatar minutes."},{"severity":"medium","title":"Small output frame","detail":"The LiveKit path renders 368x560 and center-crops your image; check it looks right on large screens."},{"severity":"low","title":"Idle timeout defaults to 60 s","detail":"idle_timeout ends idle sessions after 60 s by default; a negative value disables it, which can leave sessions billing."}],"best_for":"Character and mascot avatars (non-photoreal or stylised) from a single image.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Pricing","url":"https://lemonslice.com/pricing"},{"label":"LiveKit plugin","url":"https://docs.livekit.io/agents/models/avatar/plugins/lemonslice"},{"label":"LemonSlice LiveKit integration","url":"https://lemonslice.com/docs/self-managed/integrations/livekit-agent-integration"}],"sources":["https://lemonslice.com/pricing","https://docs.livekit.io/agents/models/avatar/plugins/lemonslice"],"confidence":"medium","unverified":"Concurrency for Creator/Professional/Scale, latency, included minutes above Starter.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":30,"concurrency":3,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":8,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.048,"custom_avatar":true,"photo_avatar":true,"byo_llm":true,"byo_tts":true,"interruptions":null,"resolution_p":560,"fps":null,"cpu_ok":null,"_notes":"Price is the vendor \"as low as\" volume rate; Starter overage $0.22, hosted mode +$0.09/min. Calls 30 min self-serve, 2 h Ultra. Over-concurrency calls billed at 2x. Free plan has no API access. LiveKit output 368x560.","_added":[]}},{"id":"runway-characters","name":"Runway Characters (GWM-1 Avatars)","vendor":"Runway","category":"avatar","summary":"Runway's realtime conversational avatars, photoreal or animated, powered by its GWM-1 world model. Usable through LiveKit, a React SDK or a hosted widget.","status":"GA","models":[{"name":"gwm1_avatars","status":"GA","notes":"Real-time conversational avatars powered by GWM-1; 24 fps."}],"transports":["LiveKit","WebRTC"],"audio":{"input":"User mic; or your LiveKit agent's audio","output":"Runway character voice, or your LiveKit TTS (which overrides the character voice)"},"languages":"Not captured.","latency":"Vendor claim (via secondary summary): 1.75 s server-side turnaround from end of user speech to character response; 24 fps.","features":["photoreal and animated characters","custom characters (avatar_id) or presets (preset_id)","LiveKit plugin (Python and Node)","React SDK","hosted widget"],"pricing":{"model":"credits","items":[{"what":"gwm1_avatars session","price":"2 credits upfront, then 2 credits per 6 seconds","unit":"per session","notes":"About $0.20 per minute"},{"what":"Credit price","price":"$0.01","unit":"per credit","notes":"Bought in the developer portal per project; sales tax may apply"}],"est_per_minute_usd":{"low":0.2,"high":0.22,"basis":"2 credits per 6 s = 20 credits/min at $0.01, plus a $0.02 upfront charge per session"},"free_tier":"None found.","source":"https://docs.dev.runwayml.com/guides/pricing/"},"limits":["Max session length: help center says 30 minutes for production API; API reference reportedly says 5 minutes (conflict)","Creating a character is free; billing starts when a realtime session starts"],"regions":"Not captured.","setup":{"steps":["Create a project and API key at dev.runwayml.com and buy credits.","Create a character (avatar_id) or pick a preset (preset_id).","Install livekit-agents[runway] and set RUNWAYML_API_SECRET.","Start runway.AvatarSession before your AgentSession; pass exactly one of avatar_id or preset_id."],"endpoint":"LiveKit plugin (livekit-agents[runway]); Runway API at docs.dev.runwayml.com","auth":"RUNWAYML_API_SECRET","snippet_lang":"python","snippet":"# uv add \"livekit-agents[runway]~=1.8\"   env: RUNWAYML_API_SECRET\nfrom livekit import agents\nfrom livekit.agents import AgentServer, AgentSession\nfrom livekit.plugins import runway\n\nserver = AgentServer()\n\n@server.rtc_session(agent_name=\"my-agent\")\nasync def my_agent(ctx: agents.JobContext):\n    session = AgentSession(\n        # stt=..., llm=..., tts=...\n    )\n    avatar = runway.AvatarSession(\n        avatar_id=\"...\",   # OR preset_id=\"...\", never both\n        max_duration=600,  # seconds\n    )\n    await avatar.start(session, room=ctx.room)\n    await session.start(\n        # room=ctx.room, agent=...\n    )"},"warnings":[{"severity":"medium","title":"Session length limits conflict","detail":"Help center lists 30 minutes for production; the API reference reportedly lists 5 minutes for gwm1_avatars. Design for reconnects and test the real cap."},{"severity":"medium","title":"Slower turn-taking than avatar-only vendors","detail":"Runway's own 1.75 s turnaround figure is slower than the sub-second claims of face-only APIs; fine for characters, less so for snappy support bots."},{"severity":"medium","title":"LiveKit TTS overrides the character","detail":"If you run TTS in LiveKit, the character's configured voice and personality are replaced."},{"severity":"low","title":"Per-session upfront charge","detail":"Each session costs 2 credits up front, so many tiny sessions cost more than the per-minute rate suggests."}],"best_for":"Creative or branded characters where visual quality matters more than sub-second latency.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"API pricing","url":"https://docs.dev.runwayml.com/guides/pricing/"},{"label":"LiveKit Runway plugin","url":"https://docs.livekit.io/agents/models/avatar/plugins/runway"},{"label":"Characters product","url":"https://runwayml.com/product/characters"}],"sources":["https://docs.dev.runwayml.com/guides/pricing/","https://docs.livekit.io/agents/models/avatar/plugins/runway","https://help.runwayml.com/hc/en-us/articles/49557780326163-Runway-Characters"],"confidence":"medium","unverified":"Max session length (conflicting), latency figure (from secondary summary of Runway pages), concurrency.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":1750,"languages":null,"max_session_min":30,"concurrency":null,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.2,"custom_avatar":true,"photo_avatar":null,"byo_llm":null,"byo_tts":true,"interruptions":null,"resolution_p":null,"fps":24,"cpu_ok":null,"_notes":"Latency is vendor server-side turnaround via a secondary summary. Session cap conflicts: 30 min (help center) vs 5 min (API reference). $0.20/min plus 2 credits upfront per session.","_added":[]}},{"id":"synthesia-interactive","name":"Synthesia Interactive Avatars","vendor":"Synthesia","category":"avatar","summary":"Synthesia avatars that lip-sync an agent's speech in realtime, available as a LiveKit Agents plugin; a session can load up to five avatars and swap between them.","status":"GA","models":[{"name":"Interactive avatars","status":"GA","notes":"Release status not labelled on the LiveKit page; requires a plan that includes interactive avatars."}],"transports":["LiveKit"],"audio":{"input":"Your LiveKit agent's TTS audio","output":"Your TTS"},"languages":"Not captured.","latency":"Not captured.","features":["swap between up to 5 avatars mid-session","Synthesia Studio avatar gallery","LiveKit plugin (Python only)"],"pricing":{"model":"subscription","items":[{"what":"Creator plan (third-party)","price":"$89","unit":"per month","notes":"Reported to include API and interactive video access; realtime per-minute rate not published"},{"what":"Enterprise","price":"Custom","unit":"","notes":""}],"est_per_minute_usd":{"low":null,"high":null,"basis":"No published realtime rate found"},"free_tier":"None found for interactive avatars.","source":"https://www.arcade.software/post/synthesia-pricing"},"limits":["Up to 5 avatar IDs per session"],"regions":"Not captured.","setup":{"steps":["Confirm your Synthesia workspace plan includes interactive avatars.","Create a workspace API key.","Copy avatar IDs from the Synthesia Studio gallery.","uv add livekit-agents[synthesia] and start the avatar session in your LiveKit agent."],"endpoint":"https://developers.synthesia.io (default SYNTHESIA_API_URL)","auth":"SYNTHESIA_API_KEY","snippet_lang":"python","snippet":"# uv add \"livekit-agents[synthesia]~=1.8\"   env: SYNTHESIA_API_KEY\nfrom livekit.plugins import synthesia\n\n# inside your agent entrypoint, after creating `session`\navatar = synthesia.AvatarSession(\n    synthesia.AvatarConfig(avatar_ids=[\"<avatar-id>\"]),  # 1 to 5 IDs\n)\nawait avatar.start(session, room=ctx.room)\n# later: swap_avatar(\"<another-id>\") or \"default\"\n# await session.start(...)"},"warnings":[{"severity":"high","title":"No public realtime price","detail":"We found no published per-minute or per-session rate for interactive avatars; you will likely need sales. Get a written rate before building."},{"severity":"medium","title":"Plan gate","detail":"The API key only works for interactive avatars if your plan includes them; lower tiers reportedly have no API access."},{"severity":"low","title":"Python only","detail":"The LiveKit plugin is Python only; no Node.js plugin listed."}],"best_for":"Enterprises already licensed on Synthesia who want the same avatars live.","open_source":false,"self_hostable":false,"compliance":"Not verified in this pass.","docs":[{"label":"LiveKit Synthesia plugin","url":"https://docs.livekit.io/agents/models/avatar/plugins/synthesia"},{"label":"Synthesia API docs","url":"https://docs.synthesia.io/"}],"sources":["https://docs.livekit.io/agents/models/avatar/plugins/synthesia","https://www.arcade.software/post/synthesia-pricing"],"confidence":"low","unverified":"Pricing, latency, concurrency, release status.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":89,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":true,"byo_tts":true,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"No public realtime rate. $89 Creator plan is a third-party figure. LiveKit plugin only (Python).","_added":[]}},{"id":"hedra-realtime","name":"Hedra Realtime Avatar","vendor":"Hedra","category":"avatar","summary":"Hedra's realtime avatar (Character-3 based) for LiveKit was sunset on 15 April 2026. Hedra's API now focuses on asynchronous image, video and audio generation jobs.","status":"Deprecated","models":[{"name":"Hedra Realtime Avatar (LiveKit plugin)","status":"Deprecated","notes":"LiveKit docs: Hedra sunset the product on 15 April 2026 and the plugin no longer functions."}],"transports":["LiveKit"],"audio":{"input":"n/a","output":"n/a"},"languages":"n/a","latency":"n/a","features":[],"pricing":{"model":"credits","items":[{"what":"Hedra API (async jobs only)","price":"Per job, prepaid USD wallet","unit":"per job","notes":"No realtime pricing exists"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Product retired"},"free_tier":"n/a","source":"https://hedra.com/docs/pages/developer/realtime-avatar/get-started"},"limits":["Product retired"],"regions":"n/a","setup":{"steps":["Do not start new builds; migrate to another LiveKit avatar plugin (Tavus, Anam, Beyond Presence, bitHuman, LemonSlice, Simli, Runway, LiveAvatar, D-ID, Synthesia)."],"endpoint":"","auth":"","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"Retired product","detail":"Hedra Realtime Avatar was sunset on 15 April 2026; the LiveKit Hedra plugin no longer works."},{"severity":"medium","title":"Stale tutorials","detail":"Hedra developer pages and blog posts may still show realtime setup steps; ignore them."},{"severity":"low","title":"Still useful offline","detail":"Hedra's async avatar/video generation API is separate and still billed per job from a prepaid wallet."}],"best_for":"Nothing realtime; listed so readers do not build on it.","open_source":false,"self_hostable":false,"compliance":"n/a","docs":[{"label":"LiveKit avatar overview","url":"https://docs.livekit.io/agents/models/avatar/"}],"sources":["https://docs.livekit.io/agents/models/avatar/plugins/hedra.md","https://hedra.com/docs/pages/developer/realtime-avatar/get-started"],"confidence":"medium","unverified":"No first-party Hedra sunset notice found; date comes from LiveKit docs (search result snippet).","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"Retired: realtime avatar sunset 15 April 2026; LiveKit plugin no longer works.","_added":[]}},{"id":"soul-machines","name":"Soul Machines","vendor":"Soul Machines (in receivership)","category":"avatar","summary":"Pioneer of autonomously animated 'digital people'. The company was placed into receivership in February 2026 with KPMG seeking a buyer; treat the platform as unavailable for new builds.","status":"Deprecated","models":[{"name":"Digital People / Studio","status":"Deprecated","notes":"Company in receivership since 5 Feb 2026 (NZ press)."}],"transports":["WebRTC"],"audio":{"input":"n/a","output":"n/a"},"languages":"n/a","latency":"n/a","features":[],"pricing":{"model":"subscription","items":[],"est_per_minute_usd":{"low":null,"high":null,"basis":"Not available"},"free_tier":"n/a","source":"https://www.nbr.co.nz/tech/soul-machines-owes-19-5m-in-receivership"},"limits":["Service continuity not guaranteed"],"regions":"n/a","setup":{"steps":["Do not start new builds; existing customers should plan a migration."],"endpoint":"","auth":"","snippet_lang":"javascript","snippet":""},"warnings":[{"severity":"high","title":"Vendor in receivership","detail":"Placed into receivership in February 2026, owing creditors over NZ$19.5 million per NBR; no buyer confirmed as of the latest reporting found (April 2026)."},{"severity":"medium","title":"Status may have changed","detail":"No reporting after April 2026 was found; check KPMG notices before assuming anything either way."},{"severity":"medium","title":"Export what you own now","detail":"Existing customers should export conversation logs, scripts and any avatar assets they have rights to while the platform is still reachable."}],"best_for":"Nothing new; listed for awareness.","open_source":false,"self_hostable":false,"compliance":"n/a","docs":[],"sources":["https://www.nbr.co.nz/tech/soul-machines-owes-19-5m-in-receivership","https://www.cbinsights.com/company/soul-machines"],"confidence":"medium","unverified":"Current status after April 2026.","cat":"video","kind":"avatar","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"Vendor in receivership since February 2026; treat as unavailable.","_added":[]}},{"id":"decart","name":"Decart Realtime (Lucy)","vendor":"Decart","category":"realtime-video","summary":"Live video-to-video editing and restyling over WebRTC: send a webcam stream plus a text prompt or reference image and get the transformed stream back. Also offers a realtime world model preview (Oasis 3).","status":"GA","models":[{"name":"lucy-2.5","status":"GA","notes":"Default realtime editing model: add/replace/remove objects, swap characters, restyle; 1280x720 landscape or portrait; fast mode at 2x price."},{"name":"lucy-restyle-2","status":"GA","notes":"Style transfer at half the price."},{"name":"lucy-vton-3.6 / 3.5","status":"GA","notes":"Realtime virtual try-on."},{"name":"oasis-3 (preview)","status":"Preview","notes":"Realtime world model; enterprise pricing available."},{"name":"lucy-2.1","status":"GA","notes":"Previous generation."}],"transports":["WebRTC"],"audio":{"input":"None (video only)","output":"None"},"languages":"Prompts in text; language support not documented.","latency":"No number on the model page; fast mode described as lower latency and higher throughput.","features":["realtime video-to-video","text prompt and reference image control","short-lived client tokens","SDKs for JS, Python, Swift, Android","auto-reconnect (5 retries)","1080p option on connect"],"pricing":{"model":"per-second","items":[{"what":"Lucy 2.5 realtime","price":"$0.02 ($0.04 fast mode)","unit":"per second of active generation","notes":"720p"},{"what":"Lucy VTON 3.6 / 3.5 realtime","price":"$0.02 ($0.04 fast for 3.5)","unit":"per second","notes":"720p"},{"what":"Lucy Restyle 2 realtime","price":"$0.01","unit":"per second","notes":"720p"},{"what":"Oasis 3 Preview realtime","price":"$0.02","unit":"per second","notes":"720p; enterprise pricing available"}],"est_per_minute_usd":{"low":0.6,"high":2.4,"basis":"$0.01/s Restyle = $0.60/min up to $0.04/s fast mode = $2.40/min; Lucy 2.5 standard $1.20/min"},"free_tier":"New accounts get free credits (amount not stated). No subscription or minimum spend.","source":"https://docs.platform.decart.ai/getting-started/pricing.md"},"limits":["Concurrent realtime sessions limited by account quota (Get Realtime Quota endpoint)","Client tokens expire 60 s after minting by default"],"regions":"Not captured; network requirements page exists.","setup":{"steps":["Create an API key on platform.decart.ai.","Add a server route that mints short-lived client tokens restricted to the models you use.","npm install @decartai/sdk in the browser app.","Capture the webcam at the model's fps/size and call client.realtime.connect with a prompt.","Close the connection when the tab is hidden to stop billing."],"endpoint":"WebRTC via @decartai/sdk (client.realtime.connect)","auth":"API key server-side; short-lived client tokens in the browser","snippet_lang":"javascript","snippet":"import { createDecartClient, models } from \"@decartai/sdk\";\n\n// Browser: token comes from your server (serverClient.tokens.create)\nconst client = createDecartClient({\n  apiKeyProvider: async () => {\n    const r = await fetch(\"/api/realtime-token\", { method: \"POST\" });\n    return (await r.json()).apiKey;\n  },\n});\n\nconst model = models.realtime(\"lucy-2.5\");\nconst cam = await navigator.mediaDevices.getUserMedia({\n  video: { frameRate: model.fps, width: model.width, height: model.height },\n});\n\nconst rt = await client.realtime.connect(cam, {\n  model,\n  onRemoteStream: (s) => { document.getElementById(\"out\").srcObject = s; },\n  initialState: { prompt: { text: \"Turn me into a claymation character\" } },\n});\n\n// Server (Express):\n// const t = await serverClient.tokens.create({ expiresIn: 300,\n//   allowedModels: [\"lucy-2.5\"] }); res.json(t);"},"warnings":[{"severity":"high","title":"Expensive per hour","detail":"$0.02/s is $72 per active hour per stream at standard speed and $144 in fast mode. Put a hard session cap and stop on visibility change."},{"severity":"medium","title":"Never ship your API key to the browser","detail":"Use server-minted client tokens with allowedModels and short expiry; the SDK refetches them on reconnect."},{"severity":"medium","title":"720p is the billed baseline","detail":"Pricing is quoted at 720p; marketing mentions 1080p. Check how 1080p is billed before enabling it."},{"severity":"medium","title":"Deepfake rules apply","detail":"Character swap and face edits on live video fall under deepfake disclosure rules (e.g. EU AI Act Article 50); label output and get consent for real people."},{"severity":"low","title":"Model names change fast","detail":"Lucy 2.1 is already 'previous generation'; read fps/size from the model config instead of hardcoding."}],"best_for":"Live AR-style filters, virtual try-on, and creative live streams.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Pricing","url":"https://docs.platform.decart.ai/getting-started/pricing.md"},{"label":"Lucy 2.5 realtime","url":"https://docs.platform.decart.ai/models/realtime/lucy-2.5.md"},{"label":"JavaScript realtime SDK","url":"https://docs.platform.decart.ai/sdks/javascript-realtime.md"},{"label":"Streaming best practices","url":"https://docs.platform.decart.ai/models/realtime/streaming-best-practices.md"}],"sources":["https://docs.platform.decart.ai/getting-started/pricing.md","https://docs.platform.decart.ai/models/realtime/lucy-2.5.md","https://docs.platform.decart.ai/sdks/javascript-realtime.md"],"confidence":"high","unverified":"Latency numbers, free credit amount, concurrency quota defaults, 1080p billing.","cat":"video","kind":"realtime-video","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.6,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":720,"fps":null,"cpu_ok":null,"_notes":"Price is Lucy Restyle 2 ($0.01/s); Lucy 2.5 is $1.20/min, fast mode $2.40/min. Billed at 720p; 1080p option exists. Free credit amount not stated.","_added":[]}},{"id":"odyssey","name":"Odyssey API (interactive world model video)","vendor":"Odyssey","category":"realtime-video","summary":"Streams an interactive AI-generated video world you can steer with text prompts in realtime, over WebRTC with JS and Python SDKs.","status":"Beta","models":[{"name":"Odyssey-2 Pro","status":"GA","notes":"Model named in the API quickstart; developer API launched 23 Jan 2026."},{"name":"Odyssey-2 Max","status":"Beta","notes":"Released April 2026 per third-party listings."},{"name":"Odyssey-3","status":"Preview","notes":"Shown on odyssey.systems as the newest foundation world model; API availability not confirmed."}],"transports":["WebRTC","WebSocket"],"audio":{"input":"None","output":"Not documented"},"languages":"Text prompts.","latency":"Third-party reports about 20-22 fps at 720p (frame every ~50 ms); no official latency figure.","features":["text-to-interactive-video","image-to-video start frames (resized to 1280x704)","interact() mid-stream prompts","broadcast/viewable streams","async simulation jobs","short-lived session JWTs for browsers"],"pricing":{"model":"subscription","items":[{"what":"Free tier","price":"$0","unit":"","notes":"Five concurrent interactive streams; total hours set by account quota"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"No public price list found; pricing page 404s"},"free_tier":"Free accounts get 5 concurrent streams; hour quota not published.","source":"https://documentation.api.odyssey.ml/session-management.md"},"limits":["Each stream max 150 s by default (start a new stream to continue)","Each connection max 60 min","Idle connection with no active stream disconnects after 15 min","Up to 10 queued simulation jobs","Total streaming hours capped by account quota"],"regions":"Not published.","setup":{"steps":["Get an API key at developer.odyssey.ml.","npm install @odysseyml/odyssey.","connect() to get a MediaStream, attach it to a video element.","startStream with a prompt, interact() to steer, endStream() and disconnect() when done.","Mint session JWTs server-side for browser use instead of exposing the key."],"endpoint":"@odysseyml/odyssey SDK (WebRTC + WebSocket signaling)","auth":"API key (ody_...) or short-lived session JWT","snippet_lang":"javascript","snippet":"import { Odyssey } from \"@odysseyml/odyssey\";\n\n// Use a server-minted session credential in production, not the raw key\nconst client = new Odyssey({ apiKey: \"ody_your_api_key_here\" });\n\nconst mediaStream = await client.connect();\ndocument.querySelector(\"video\").srcObject = mediaStream;\n\nawait client.startStream({ prompt: \"A cat on a sunny windowsill\" });\nawait client.interact({ prompt: \"Pet the cat\" });\n// each stream lasts max 150 s by default\nawait client.endStream();\nclient.disconnect();"},"warnings":[{"severity":"high","title":"No published pricing","detail":"Odyssey's pricing URL redirects to a 404 on odyssey.systems; usage is metered in hours against an account quota. Get rates in writing before launch."},{"severity":"medium","title":"2.5 minute stream cap","detail":"Streams end after 150 s by default; the docs suggest a dual-session rollover pattern that uses two concurrent slots."},{"severity":"medium","title":"Research-grade output","detail":"World models drift and lose consistency over time; do not promise game-like persistence."},{"severity":"low","title":"Domain move","detail":"The company appears to be moving from odyssey.ml to odyssey.systems; docs still live at documentation.api.odyssey.ml."}],"best_for":"Interactive storytelling, prototypes and demos of playable AI video.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"API quick start","url":"https://documentation.api.odyssey.ml/api-quick-start.md"},{"label":"Stream duration limits","url":"https://documentation.api.odyssey.ml/stream-duration-limits.md"},{"label":"Session management","url":"https://documentation.api.odyssey.ml/session-management.md"}],"sources":["https://documentation.api.odyssey.ml/session-management.md","https://documentation.api.odyssey.ml/stream-duration-limits.md","https://documentation.api.odyssey.ml/api-quick-start.md","https://odyssey.ml/the-gpt-2-moment-for-world-models","https://odyssey.systems/"],"confidence":"medium","unverified":"Pricing, fps/latency, which model the API serves today, Odyssey-3 API availability.","cat":"video","kind":"realtime-video","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":2.5,"concurrency":5,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":720,"fps":20,"cpu_ok":null,"_notes":"Streams end after 150 s by default; connections max 60 min. Concurrency 5 is the free tier. No published pricing. Resolution and fps (20-22) are third-party reports.","_added":[]}},{"id":"fal-realtime","name":"fal Realtime endpoints","vendor":"fal","category":"realtime-video","summary":"WebSocket realtime inference on fal: a persistent connection to a warm runner for fast image-to-image loops (e.g. LCM / SDXL Turbo), plus serverless GPUs where you can deploy your own WebRTC realtime video or world model app.","status":"GA","models":[{"name":"fal-ai/fast-lcm-diffusion","status":"GA","notes":"SDXL with Latent Consistency Models; realtime endpoint."},{"name":"fal-ai/fast-turbo-diffusion","status":"GA","notes":"Optimised SDXL Turbo; realtime endpoint."},{"name":"Custom realtime apps (e.g. Matrix-Game world model demo)","status":"Beta","notes":"Deploy with fal deploy; /webrtc endpoint; fal Serverless deploy access is approved per account."}],"transports":["WebSocket","WebRTC"],"audio":{"input":"None for image endpoints","output":"None"},"languages":"Text prompts.","latency":"Vendor: realtime requests skip the queue and reuse a warm runner; first connection can still cold start.","features":["persistent WebSocket with msgpack","token provider with short-lived JWTs","proxy URL pattern","deploy your own realtime WebRTC apps on serverless GPUs","server errors are never billed"],"pricing":{"model":"per-second","items":[{"what":"Model API realtime endpoints","price":"Per model (unit varies)","unit":"see model page","notes":"fast-lcm-diffusion page showed '$0 per compute second' to our fetcher, which looks like a rendering artefact; verify"},{"what":"H100 serverless GPU","price":"$4.50 list, as low as $2.49","unit":"per hour","notes":"For your own deployed realtime apps"},{"what":"H200 serverless GPU","price":"$6.00 list, as low as $2.99","unit":"per hour","notes":""},{"what":"B200 serverless GPU","price":"$7.99 list, as low as $5.49","unit":"per hour","notes":""}],"est_per_minute_usd":{"low":0.04,"high":0.075,"basis":"One H100 runner for a self-deployed realtime app: $2.49-$4.50/h divided by 60 (one stream per GPU assumed)"},"free_tier":"Not verified.","source":"https://fal.ai/pricing"},"limits":["Only models with an explicit /realtime endpoint work with fal.realtime.connect","Cold start on first connection","Serverless deploy access needs approval"],"regions":"Not captured.","setup":{"steps":["Create a fal API key.","Add a server-side proxy route or token provider so the browser never sees the key.","npm install @fal-ai/client.","Open fal.realtime.connect to a realtime-capable model and send frames/prompts in a loop.","For live video, deploy your own app with a /webrtc endpoint (see the realtime world model example)."],"endpoint":"fal.realtime.connect(<model>) or wss://ws.fal.run/{model_id}","auth":"FAL_KEY via server proxy or short-lived JWT token provider","snippet_lang":"javascript","snippet":"import { fal } from \"@fal-ai/client\";\n\n// Route through your server so the key stays secret\nfal.config({ proxyUrl: \"/api/fal/proxy\" });\n\nconst connection = fal.realtime.connect(\"fal-ai/fast-lcm-diffusion\", {\n  onResult: (result) => {\n    // result shape: check the model's realtime schema\n    console.log(result);\n  },\n  onError: (err) => console.error(err),\n});\n\n// Send a new frame or prompt whenever it changes\nconnection.send({\n  prompt: \"a watercolor fox\",\n  image_url: canvas.toDataURL(\"image/jpeg\", 0.7),\n  strength: 0.6,\n});\n// connection.close() when done"},"warnings":[{"severity":"medium","title":"Image loop, not true video","detail":"The hosted realtime endpoints are image-to-image models you call per frame; you build the frame loop, throttling and temporal consistency yourself."},{"severity":"medium","title":"Cold starts","detail":"The first connection can hit a cold start; keep the socket open and warm the runner before users arrive."},{"severity":"medium","title":"Self-deployed GPUs bill while connected","detail":"A custom realtime app holds a GPU for the whole session; one stream per H100 is $2.49-$4.50 per hour."},{"severity":"low","title":"Realtime prices not on the main pricing page","detail":"Per-model realtime prices were not captured; check each model page or the pricing API before budgeting."}],"best_for":"Prototyping realtime image transformation loops and hosting your own realtime video models on demand GPUs.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Realtime model APIs","url":"https://fal.ai/docs/model-apis/real-time"},{"label":"Deploy a realtime world model","url":"https://fal.ai/docs/examples/video-generation/deploy-realtime-world-model"},{"label":"Pricing","url":"https://fal.ai/pricing"}],"sources":["https://fal.ai/docs/model-apis/real-time","https://fal.ai/pricing","https://fal.ai/models/fal-ai/fast-lcm-diffusion","https://fal.ai/docs/examples/video-generation/deploy-realtime-world-model"],"confidence":"medium","unverified":"Per-model realtime prices, result payload shape, free credits, Krea/H3 Max realtime pricing on fal (secondary sources only: $0.08/s after a promo).","cat":"video","kind":"realtime-video","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.04,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"Price assumes one self-deployed stream per H100 at $2.49/h (list $4.50/h). Hosted realtime endpoints are per-frame image models billed per model.","_added":[]}},{"id":"daydream-streamdiffusion","name":"Daydream (hosted StreamDiffusion)","vendor":"Daydream / Livepeer","category":"realtime-video","summary":"Hosted StreamDiffusion API for live stream restyling on the Livepeer GPU network, plus the open-source Scope toolkit (StreamDiffusionV2, LongLive, Krea Realtime 14B). StreamDiffusion itself is open source and self-hostable.","status":"Beta","models":[{"name":"StreamDiffusion (SD/SDXL)","status":"Beta","notes":"Hosted API; SDXL added Nov 2025."},{"name":"Scope toolkit","status":"Preview","notes":"Open-source local toolkit in community alpha."}],"transports":["WebRTC"],"audio":{"input":"None","output":"None"},"languages":"Text prompts.","latency":"Vendor-described sub-second latency (secondary source).","features":["live video restyling","create/update stream API","Livepeer decentralised GPU network","open-source self-host path"],"pricing":{"model":"credits","items":[{"what":"daydream.live/pricing plans","price":"$0 / $10 / $30","unit":"per month","notes":"50 / 500 / 1,750 credits; play meters at 1.25 credits per minute. This page appears to describe a consumer 'play' product that streams audio, so it may not apply to the StreamDiffusion API."}],"est_per_minute_usd":{"low":null,"high":null,"basis":"API pricing not confirmed"},"free_tier":"Unclear for the API.","source":"https://daydream.live/pricing"},"limits":["'Submit StreamDiffusion Prompt' endpoint marked deprecated in the API reference"],"regions":"Decentralised Livepeer orchestrators.","setup":{"steps":["Get a Daydream API key.","Use the Daydream docs as the source of truth for the current base URL and create-stream route.","Publish your camera to the stream and play back the transformed output.","Or self-host StreamDiffusion / Scope on your own GPU."],"endpoint":"See Daydream API reference (base URL not verified)","auth":"API key","snippet_lang":"javascript","snippet":""},"warnings":[{"severity":"high","title":"Pricing and product direction unclear","detail":"Daydream's pricing page now describes a credits-based product that 'streams audio' and trains 'Styles', which does not obviously match the StreamDiffusion video API. Confirm before building."},{"severity":"medium","title":"Deprecated prompt endpoint","detail":"The old 'Submit StreamDiffusion Prompt' route is deprecated; use the update-stream endpoint."},{"severity":"medium","title":"Decentralised network variability","detail":"Jobs route to independent orchestrators, so latency and capacity can vary."}],"best_for":"VJ, creative live streams, and teams willing to self-host StreamDiffusion.","open_source":true,"self_hostable":true,"compliance":"Not verified.","docs":[{"label":"Livepeer Daydream overview","url":"https://docs.livepeer.org/v2/solutions/Daydream/overview"},{"label":"Daydream gateway","url":"https://docs.livepeer.org/v2/gateways/using-gateways/gateway-providers/daydream-gateway"}],"sources":["https://docs.livepeer.org/v2/solutions/Daydream/overview","https://www.businesswire.com/news/home/20251106860538/en/Daydream-Launches-Scope-and-Expands-StreamDiffusion-with-SDXL-Support-Advancing-the-Open-Source-Real-Time-AI-Video-Ecosystem","https://daydream.live/pricing"],"confidence":"low","unverified":"API pricing, base URL, current status of hosted API, latency.","cat":"video","kind":"realtime-video","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":true,"self_hostable":true,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"API pricing unconfirmed; pricing page ($0/$10/$30 credit plans) may describe a different consumer product. Vendor says sub-second latency without a number.","_added":[]}},{"id":"google-genie","name":"Google Project Genie (Genie 3)","vendor":"Google DeepMind","category":"realtime-video","summary":"Genie 3 generates explorable interactive worlds in realtime. As of the latest reporting found it is a research prototype for Google AI Ultra subscribers, with no public developer API.","status":"Preview","models":[{"name":"Genie 3 / Project Genie","status":"Preview","notes":"Launched to US AI Ultra subscribers around 30 Jan 2026; Google said it is investigating an API."}],"transports":[],"audio":{"input":"n/a","output":"n/a"},"languages":"n/a","latency":"n/a","features":["interactive world generation (consumer app only)"],"pricing":{"model":"subscription","items":[{"what":"Google AI Ultra (consumer)","price":"$250","unit":"per month","notes":"US; not an API"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"No API"},"free_tier":"None","source":"https://itdaily.com/news/cloud/google-launches-project-genie"},"limits":["No API"],"regions":"US first at launch.","setup":{"steps":["No developer setup possible; watch Google AI for Developers for an API announcement."],"endpoint":"","auth":"","snippet_lang":"python","snippet":""},"warnings":[{"severity":"high","title":"No API","detail":"Do not plan a product on Genie; there is no developer API as of the reporting found."},{"severity":"low","title":"Reporting may be stale","detail":"Sources are from early 2026; check Google DeepMind's blog for newer access."},{"severity":"low","title":"Buildable alternatives","detail":"For a world model you can call today, look at Odyssey's API or Decart's Oasis 3 preview (realtime, $0.02/s at 720p)."}],"best_for":"Watching the space; not buildable.","open_source":false,"self_hostable":false,"compliance":"n/a","docs":[],"sources":["https://itdaily.com/news/cloud/google-launches-project-genie","https://yourstory.com/ai-story/google-deepmind-project-genie-launch"],"confidence":"medium","unverified":"Any change after early 2026.","cat":"video","kind":"realtime-video","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"No developer API; consumer access only via Google AI Ultra ($250/month, US).","_added":[]}},{"id":"gemini-live-video","name":"Gemini Live API (video input)","vendor":"Google","category":"vision","summary":"Pointer entry: the Gemini Live API accepts camera or screen frames alongside audio. This entry covers only what video frames cost; see the voice segment for the rest.","status":"Preview","models":[{"name":"gemini-3.8-live","status":"GA","notes":"Image/video input $1.00 per 1M tokens or $0.002/min."},{"name":"gemini-3.1-flash-live-preview","status":"Preview","notes":"Same price card as 3.8 Live."},{"name":"gemini-2.5-flash-native-audio-preview-12-2025","status":"Preview","notes":"Audio/video input $3.00 per 1M tokens."},{"name":"gemini-robotics-er-2-streaming-preview","status":"Preview","notes":"$1.00/M input incl. video through 31 Dec 2026, $2.00 from 1 Jan 2027."}],"transports":["WebSocket"],"audio":{"input":"PCM audio","output":"Native audio"},"languages":"See voice segment.","latency":"See voice segment.","features":["camera and screen frames as JPEG/PNG","mediaResolution control of tokens per frame","session resumption / compression for longer sessions"],"pricing":{"model":"per-minute","items":[{"what":"Video/image input, Gemini 3.8 Live","price":"$1.00 per 1M tokens or $0.002","unit":"per minute","notes":"Paid tier"},{"what":"Tokens per video frame (Gemini 3)","price":"70 tokens (low/medium/default), 280 (high)","unit":"per frame","notes":"Max 1 frame per second in Live API"},{"what":"Audio for comparison, 3.8 Live","price":"$0.005 in / $0.018 out","unit":"per minute","notes":""}],"est_per_minute_usd":{"low":0.002,"high":0.0042,"basis":"Video input only. Low = Google's listed $0.002/min; high = our arithmetic at 1 fps x 70 tokens x 60 s = 4,200 tokens at $1/M. Audio is billed on top."},"free_tier":"See Gemini API free tier (not checked here).","source":"https://ai.google.dev/gemini-api/docs/pricing"},"limits":["Max 1 frame per second","Audio plus video sessions limited to 2 minutes unless you use session management/compression","Context window 128k tokens for native audio models"],"regions":"See voice segment.","setup":{"steps":["Get a Gemini API key (or use ephemeral tokens in browsers).","Open a Live session with response_modalities AUDIO and media_resolution LOW.","Send JPEG frames at 1 fps or less with send_realtime_input(video=...).","Enable context window compression for sessions over 2 minutes."],"endpoint":"Live API WebSocket via google-genai SDK","auth":"GEMINI_API_KEY or ephemeral token","snippet_lang":"python","snippet":"import asyncio\nfrom google import genai\nfrom google.genai import types\n\nclient = genai.Client()  # GEMINI_API_KEY\nconfig = {\n    \"response_modalities\": [\"AUDIO\"],\n    \"media_resolution\": \"MEDIA_RESOLUTION_LOW\",  # 70 tokens/frame on Gemini 3\n}\n\nasync def main(jpeg_frames):\n    async with client.aio.live.connect(model=\"gemini-3.8-live\", config=config) as s:\n        for jpg in jpeg_frames:            # at most 1 frame per second\n            await s.send_realtime_input(\n                video=types.Blob(data=jpg, mime_type=\"image/jpeg\"))\n            await asyncio.sleep(1)\n        await s.send_realtime_input(text=\"What do you see?\")\n        async for msg in s.receive():\n            pass  # handle audio chunks / transcripts\n\n# asyncio.run(main(frames))"},"warnings":[{"severity":"high","title":"2 minute cap with video","detail":"Audio plus video sessions are limited to 2 minutes unless you configure session management; plan compression or resumption."},{"severity":"medium","title":"1 fps maximum","detail":"The model sees at most one frame per second; it will miss fast motion. Not suitable for frame-accurate tasks."},{"severity":"medium","title":"Set media resolution low","detail":"Video frames default to 70 tokens on Gemini 3 but high is 280; images at high or ultra_high cost far more."},{"severity":"low","title":"Model IDs churn","detail":"Live models move fast (3.1 Flash Live preview, 3.8 Live, 2.5 native audio preview); pin a model and watch deprecations."}],"best_for":"Cheapest realtime 'see and talk' assistant; video input adds only fractions of a cent per minute.","open_source":false,"self_hostable":false,"compliance":"Google Cloud / Gemini API terms; see voice segment.","docs":[{"label":"Live API capabilities","url":"https://ai.google.dev/gemini-api/docs/live-api/capabilities"},{"label":"Media resolution","url":"https://ai.google.dev/gemini-api/docs/media-resolution"},{"label":"Pricing","url":"https://ai.google.dev/gemini-api/docs/pricing"}],"sources":["https://ai.google.dev/gemini-api/docs/pricing","https://ai.google.dev/gemini-api/docs/media-resolution","https://ai.google.dev/gemini-api/docs/live-api/capabilities"],"confidence":"high","unverified":"Whether gemini-3.8-live is GA or preview label; per-frame tokens for 2.5 models.","cat":"video","kind":"vision","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":2,"concurrency":null,"free_tier":null,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":null,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":0.002,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":1,"cpu_ok":null,"_notes":"Price is video input only on Gemini 3.8 Live (audio billed on top). Audio plus video sessions capped at 2 min without session management. fps is max input frame rate.","_added":[]}},{"id":"openai-realtime-image","name":"OpenAI Realtime API (image input)","vendor":"OpenAI","category":"vision","summary":"Pointer entry: OpenAI's realtime speech models accept still images in the conversation. There is no continuous video track; apps send sampled frames (often 1 fps) as input_image items.","status":"GA","models":[{"name":"gpt-realtime-2.1","status":"GA","notes":"Image input $5.00 / cached $0.50 per 1M tokens."},{"name":"gpt-realtime-2.1-mini","status":"GA","notes":"Image input $0.80 / cached $0.08 per 1M tokens."},{"name":"gpt-realtime-2, gpt-realtime-1.5, gpt-realtime","status":"GA","notes":"Same price card as 2.1."},{"name":"gpt-realtime-mini","status":"GA","notes":"Same as 2.1-mini."},{"name":"gpt-live-1","status":"GA","notes":"Listed as GPT-Live sessions at $0.05/min; image/video support not confirmed."}],"transports":["WebRTC","WebSocket"],"audio":{"input":"Speech","output":"Speech"},"languages":"See voice segment.","latency":"See voice segment.","features":["input_image items with base64 data URL","text + image in one message","frame sampling done by your app (LiveKit samples 1 fps by default)"],"pricing":{"model":"per-minute","items":[{"what":"Image input, gpt-realtime-2.1","price":"$5.00","unit":"per 1M tokens","notes":"Cached $0.50"},{"what":"Image input, gpt-realtime-2.1-mini","price":"$0.80","unit":"per 1M tokens","notes":"Cached $0.08"},{"what":"gpt-live-1 session","price":"$0.05","unit":"per minute","notes":"Backend model and tool use billed separately"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Depends on tokens per image (detail level, size) which we did not verify for realtime models"},"free_tier":"None.","source":"https://developers.openai.com/api/docs/pricing"},"limits":["No native video stream; you choose frame rate and size","Large base64 images over WebRTC data channel reported as troublesome by developers"],"regions":"See voice segment.","setup":{"steps":["Open a Realtime session (WebRTC in browser with an ephemeral key, or WebSocket server-side).","Capture a frame from the camera or screen to a small JPEG (e.g. 512 px).","Send conversation.item.create with input_text + input_image, then response.create.","Throttle frames; only send on change or on user request."],"endpoint":"wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1","auth":"Bearer API key server-side; ephemeral client secret in browser","snippet_lang":"javascript","snippet":"// ws: an open Realtime WebSocket (or RTCDataChannel \"oai-events\")\nfunction sendFrame(ws, base64Jpeg, question) {\n  ws.send(JSON.stringify({\n    type: \"conversation.item.create\",\n    item: {\n      type: \"message\",\n      role: \"user\",\n      content: [\n        { type: \"input_text\", text: question },\n        { type: \"input_image\", image_url: `data:image/jpeg;base64,${base64Jpeg}` },\n      ],\n    },\n  }));\n  ws.send(JSON.stringify({ type: \"response.create\" }));\n}\n\n// Grab a frame: draw <video> to a 512px canvas, then\n// canvas.toDataURL(\"image/jpeg\", 0.7).split(\",\")[1]"},"warnings":[{"severity":"medium","title":"Images stay in context and cost every turn","detail":"Each image becomes conversation tokens that are re-read on later turns; prune old frames or you pay for them repeatedly (cached rate helps)."},{"severity":"medium","title":"Not real video understanding","detail":"Sampled stills miss motion; for continuous monitoring use a dedicated video service."},{"severity":"medium","title":"Prefer WebSocket for big images","detail":"Developers reported trouble sending large base64 images over the WebRTC data channel; downscale to around 512 px."},{"severity":"low","title":"gpt-live-1 vision unclear","detail":"We could not confirm that gpt-live-1 accepts images; test before choosing it for vision."}],"best_for":"Voice assistants that occasionally need to look at a screenshot or camera frame.","open_source":false,"self_hostable":false,"compliance":"OpenAI API terms; see voice segment.","docs":[{"label":"Pricing","url":"https://developers.openai.com/api/docs/pricing"},{"label":"Realtime client events","url":"https://developers.openai.com/docs/api-reference/realtime-client-events"},{"label":"LiveKit OpenAI realtime (video sampling)","url":"https://docs.livekit.io/agents/models/realtime/openai/"}],"sources":["https://developers.openai.com/api/docs/pricing","https://community.openai.com/t/realtime-model-image-input/1355688","https://docs.livekit.io/agents/models/realtime/openai/"],"confidence":"medium","unverified":"Tokens per image for realtime models, gpt-live-1 vision support.","cat":"video","kind":"vision","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":true,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"Still images only (no video track), billed as image tokens ($5/M on gpt-realtime-2.1, $0.80/M on mini). Per-minute cost depends on frame rate and size.","_added":[]}},{"id":"overshoot","name":"Overshoot","vendor":"Overshoot (YC W26)","category":"vision","summary":"Realtime vision API: publish a live camera to a stream over LiveKit, then ask OpenAI-style chat completions about the latest frames using fast open VLMs it hosts, or pass through to Gemini, Claude or GPT.","status":"Beta","models":[{"name":"google/gemma-4-26B-A4B-it","status":"GA","notes":"Overshoot-hosted; $0.06/M input, $0.33/M output."},{"name":"google/gemma-4-31B-it","status":"GA","notes":"$0.12/M input, $0.36/M output."},{"name":"Qwen/Qwen3.6-27B-FP8","status":"GA","notes":"$0.29/M input, $2.40/M output."},{"name":"Qwen/Qwen3.6-35B-A3B-FP8","status":"GA","notes":"$0.16/M input, $1.10/M output."},{"name":"Hcompany/Holo3-35B-A3B and Holo-3.1","status":"GA","notes":"$0.25/M input, $1.80/M output."},{"name":"Passthrough: Gemini 3.x, Claude 4.x, GPT-5.4 family","status":"GA","notes":"Upstream latency is seconds, not sub-second."}],"transports":["LiveKit","WebRTC"],"audio":{"input":"None","output":"None (text)"},"languages":"Depends on chosen model.","latency":"Vendor: hosted models sized for sub-second time-to-first-token on single-frame inputs; no measured figure.","features":["live stream frame references (ovs://streams/{id}?frame_index=-1)","OpenAI-compatible chat completions","open VLMs and proprietary passthrough","public pricing endpoint","prepaid credits from $1"],"pricing":{"model":"credits","items":[{"what":"gemma-4-26B-A4B (hosted)","price":"$0.06 in / $0.33 out","unit":"per 1M tokens","notes":"6 / 33 microcents per token from GET /billing/pricing"},{"what":"Qwen3.6-27B-FP8 (hosted)","price":"$0.29 in / $2.40 out","unit":"per 1M tokens","notes":""},{"what":"gemini-3-flash-preview passthrough","price":"$0.50 in / $3.00 out","unit":"per 1M tokens","notes":""},{"what":"claude-sonnet-4-6 passthrough","price":"$3.00 in / $15.00 out","unit":"per 1M tokens","notes":""},{"what":"Stream time","price":"Not shown","unit":"","notes":"No per-minute stream charge in the public price list; verify"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Token-based; depends on frames per query and tokens per frame (not published)"},"free_tier":"None found; prepaid credits, minimum $1.","source":"https://api.overshoot.ai/billing/pricing"},"limits":["Frames retained for 600 seconds","Frames may be compressed or resized before inference","A 'ready' model can still return 503"],"regions":"Not published.","setup":{"steps":["Buy prepaid credits and create an API key.","POST /v1beta/streams to create a stream; get publish.url and publish.token.","Publish your webcam to that LiveKit room.","POST /v1beta/chat/completions referencing ovs://streams/{id}?frame_index=-1 for the latest frame."],"endpoint":"https://api.overshoot.ai/v1beta","auth":"Authorization: Bearer <OVERSHOOT_API_KEY>","snippet_lang":"javascript","snippet":"const H = {\n  Authorization: `Bearer ${process.env.OVERSHOOT_API_KEY}`,\n  \"Content-Type\": \"application/json\",\n};\nconst API = \"https://api.overshoot.ai/v1beta\";\n\n// 1) create a stream, then publish the camera to publish.url\n//    with publish.token using the LiveKit client SDK\nconst stream = await (await fetch(`${API}/streams`, { method: \"POST\", headers: H })).json();\n\n// 2) ask about the newest frame\nconst r = await fetch(`${API}/chat/completions`, {\n  method: \"POST\",\n  headers: H,\n  body: JSON.stringify({\n    model: \"google/gemma-4-26B-A4B-it\",\n    messages: [{ role: \"user\", content: [\n      { type: \"text\", text: \"Is anyone at the door?\" },\n      { type: \"image_url\", image_url: { url: `ovs://streams/${stream.id}?frame_index=-1` } },\n    ]}],\n  }),\n});\nconsole.log((await r.json()).choices[0].message.content);"},"warnings":[{"severity":"medium","title":"Young startup, beta API","detail":"v1beta paths and a YC W26 company; expect breaking changes and keep an abstraction layer."},{"severity":"medium","title":"Passthrough is not realtime","detail":"Claude/GPT/Gemini passthrough adds seconds of latency; use hosted Gemma/Qwen for sub-second loops."},{"severity":"medium","title":"You poll, it does not push","detail":"Analysis happens per chat completion you send; continuous monitoring means calling it on a timer, which multiplies token cost."},{"severity":"low","title":"Short frame history","detail":"Only the last 600 seconds of frames are addressable."}],"best_for":"Cheap, fast 'what is happening on camera right now' questions with open VLMs.","open_source":false,"self_hostable":false,"compliance":"Not verified.","docs":[{"label":"Quickstart","url":"https://docs.overshoot.ai/quickstart.md"},{"label":"Models","url":"https://docs.overshoot.ai/models.md"},{"label":"List pricing","url":"https://docs.overshoot.ai/api-reference/list-pricing.md"},{"label":"Limits and retention","url":"https://docs.overshoot.ai/api-reference/limits.md"}],"sources":["https://api.overshoot.ai/billing/pricing","https://docs.overshoot.ai/models.md","https://docs.overshoot.ai/quickstart.md","https://www.ycombinator.com/companies/overshoot"],"confidence":"medium","unverified":"Stream-time charges, concurrency, tokens per frame, latency.","cat":"video","kind":"vision","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":false,"free_credit_usd":null,"entry_plan_usd_month":0,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":false,"self_hostable":false,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":null,"_notes":"Token-billed prepaid credits from $1; no stream-time charge shown. Vendor targets sub-second TTFT on hosted models without a number. Frames kept 600 s.","_added":[]}},{"id":"roboflow-webrtc","name":"Roboflow Serverless Video Streaming (WebRTC)","vendor":"Roboflow","category":"vision","summary":"Stream a webcam, RTSP camera or file to Roboflow Cloud over WebRTC and run detection models or Workflows on every frame, getting annotated video and JSON results back.","status":"Beta","models":[{"name":"webrtc-gpu-medium","status":"GA","notes":"Default plan, recommended for most Workflows; 60 min per credit."},{"name":"webrtc-gpu-small","status":"GA","notes":"Lower cost; 80 min per credit."},{"name":"webrtc-gpu-large","status":"GA","notes":"Required for SAM3 and Rapid Models; about 5 FPS; 30 min per credit."},{"name":"CPU video streams","status":"GA","notes":"10 hours per credit."}],"transports":["WebRTC"],"audio":{"input":"None","output":"None"},"languages":"n/a","latency":"Quality and FPS can take up to a minute to ramp up at 1080p30 (vendor docs).","features":["any Roboflow model or Workflow","bidirectional video track with annotations","reliable JSON data channel","regions us / eu / ap","same SDK works against self-hosted Inference server"],"pricing":{"model":"credits","items":[{"what":"GPU small / medium / large","price":"1 credit per 80 / 60 / 30 min","unit":"of video","notes":"Billed hourly by plan"},{"what":"CPU stream","price":"1 credit per 10 h","unit":"of video","notes":""},{"what":"Credit price (Core plan)","price":"$2.86 to $3.90","unit":"per credit","notes":"Monthly packs; on-demand listed at $6, prepaid from $4"},{"what":"Core plan","price":"$39","unit":"per month","notes":"Free tier: 10 credits/month"}],"est_per_minute_usd":{"low":0.005,"high":0.2,"basis":"CPU 0.1 credit/h at $2.86 = $0.005/min; GPU large 2 credits/h at $6 = $0.20/min; GPU medium $0.05-$0.10/min"},"free_tier":"Free plan includes 10 credits per month (about 10 hours of GPU-medium streaming).","source":"https://roboflow.com/credits"},"limits":["10 concurrent streams per workspace by default","Billing starts when the serverless function spawns and WebRTC connects","Streams stop with HTTP 402 when credits run out"],"regions":"us, eu, ap","setup":{"steps":["Create a Roboflow workspace, API key and a model or Workflow.","pip install \"inference-sdk[webrtc]\" (or npm install @roboflow/inference-sdk).","Point the client at https://serverless.roboflow.com and start a WebRTC stream with a StreamConfig (plan, region, outputs).","Handle JSON predictions from the data channel; stop the session to stop billing."],"endpoint":"https://serverless.roboflow.com (WebRTC signalling via SDK)","auth":"Roboflow API key","snippet_lang":"python","snippet":"# pip install \"inference-sdk[webrtc]\"  (SDK marked experimental)\n# Sketch based on documented parameters; check field names\n# against the examples/webrtc_sdk scripts in roboflow/inference.\nfrom inference_sdk import InferenceHTTPClient\nfrom inference_sdk.webrtc import StreamConfig, WebcamSource\n\nclient = InferenceHTTPClient(\n    api_url=\"https://serverless.roboflow.com\",\n    api_key=\"YOUR_API_KEY\",\n)\n\nsession = client.webrtc.stream(\n    source=WebcamSource(),\n    workflow=\"your-workflow-id\",\n    workspace=\"your-workspace\",\n    config=StreamConfig(\n        data_output=[\"predictions\"],\n        requested_plan=\"webrtc-gpu-small\",\n        requested_region=\"us\",\n    ),\n)\n\n@session.on_data(\"predictions\")\ndef handle(predictions, metadata):\n    print(metadata.frame_id, predictions)\n\nsession.run()"},"warnings":[{"severity":"high","title":"Streams die at zero credits","detail":"A user reported HTTP 402 CreditsExceededError mid-stream even with ~50 credits showing; keep a buffer and enable top-ups for production cameras."},{"severity":"medium","title":"Billed from connection, not from first result","detail":"Billing starts when the function spawns and WebRTC connects, including the up to one minute ramp-up."},{"severity":"medium","title":"Credit price varies a lot","detail":"The same credit costs $2.86 to $6 depending on pack versus on-demand; budget at the on-demand rate."},{"severity":"low","title":"Experimental SDK","detail":"The WebRTC SDK is labelled experimental; pin versions."}],"best_for":"Object detection, counting and safety monitoring on live cameras without running your own GPU.","open_source":true,"self_hostable":true,"compliance":"Not verified in this pass.","docs":[{"label":"Serverless video streaming","url":"https://docs.roboflow.com/deploy/serverless-video-streaming-api"},{"label":"WebRTC SDK reference","url":"https://inference.roboflow.com/reference/inference_sdk/webrtc/config"},{"label":"Credits","url":"https://roboflow.com/credits"},{"label":"Pricing","url":"https://roboflow.com/pricing"}],"sources":["https://docs.roboflow.com/deploy/serverless-video-streaming-api","https://roboflow.com/credits","https://roboflow.com/pricing","https://discuss.roboflow.com/t/http-402-payment-required-creditsexceedederror-when-using-webrtc-gpu-medium-on-research-plan/11853"],"confidence":"medium","unverified":"Exact SDK call signature, max stream duration, when billing stops.","cat":"video","kind":"vision","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":10,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":39,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":true,"websocket":null,"sip":null,"open_source":true,"self_hostable":true,"price_per_min":0.005,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":true,"_notes":"Price is CPU streaming; GPU medium about $0.05-$0.10/min, GPU large up to $0.20/min. Free plan 10 credits/month. 10 concurrent streams per workspace by default. GPU large runs about 5 FPS.","_added":[]}},{"id":"nvidia-vss","name":"NVIDIA VSS blueprint / Cosmos Reason","vendor":"NVIDIA","category":"vision","summary":"NVIDIA's Video Search and Summarization blueprint runs VLMs (Cosmos Reason, Qwen3-VL) over live RTSP streams for captions and alerts. It is self-hosted; build.nvidia.com offers free preview API calls, not a hosted streaming endpoint.","status":"Preview","models":[{"name":"cosmos-reason2-8b","status":"GA","notes":"Downloadable NIM; 256K input context."},{"name":"VSS 3.2 Real-Time VLM component","status":"GA","notes":"Self-hosted; processes stream chunks at a user-defined interval."}],"transports":["RTSP"],"audio":{"input":"n/a","output":"n/a"},"languages":"n/a","latency":"Chunk-based (user-defined chunk duration); not frame-level realtime.","features":["live RTSP stream alerts","captioning and anomaly detection","self-hosted NIM microservices"],"pricing":{"model":"free","items":[{"what":"build.nvidia.com preview APIs","price":"Up to 5,000 free credits","unit":"per new account","notes":"Trial only"},{"what":"Production self-hosting","price":"NVIDIA AI Enterprise licence","unit":"","notes":"Plus your own GPUs"}],"est_per_minute_usd":{"low":null,"high":null,"basis":"Self-hosted; depends on your GPUs and licence"},"free_tier":"Preview API credits for new accounts.","source":"https://www.nvidia.com/en-us/use-cases/video-analytics-ai-agents/"},"limits":["No hosted realtime streaming endpoint confirmed"],"regions":"Wherever you deploy.","setup":{"steps":["Deploy the VSS blueprint on your GPUs (see docs.nvidia.com/vss).","Register RTSP streams and configure chunk duration and alert prompts.","Use build.nvidia.com preview APIs only for trials."],"endpoint":"Self-hosted VSS REST API","auth":"NGC / NVIDIA API key","snippet_lang":"python","snippet":""},"warnings":[{"severity":"medium","title":"Not a hosted realtime API","detail":"Plan on running GPUs yourself; the hosted catalog is for trials."},{"severity":"medium","title":"Licence needed for production","detail":"Downloadable NIMs need an NVIDIA AI Enterprise licence in production."},{"severity":"low","title":"Chunked, not per-frame","detail":"Alerts run on segments at an interval you set, so reaction time is at least one chunk."}],"best_for":"Enterprises with GPUs building camera monitoring and video search.","open_source":false,"self_hostable":true,"compliance":"Depends on your deployment.","docs":[{"label":"VSS docs","url":"https://docs.nvidia.com/vss/3.2.1/"},{"label":"VSS blueprint","url":"https://build.nvidia.com/nvidia/video-search-and-summarization/blueprintcard"},{"label":"Cosmos Reason 2 8B","url":"https://build.nvidia.com/nvidia/cosmos-reason2-8b/modelcard"}],"sources":["https://docs.nvidia.com/vss/3.2.1/","https://www.nvidia.com/en-us/use-cases/video-analytics-ai-agents/","https://build.nvidia.com/nvidia/cosmos-reason2-8b/modelcard"],"confidence":"low","unverified":"Whether any hosted streaming endpoint exists; Cosmos 3 status.","cat":"video","kind":"vision","verified_at":"2026-10-10","facts":{"latency_ms":null,"languages":null,"max_session_min":null,"concurrency":null,"free_tier":true,"free_credit_usd":null,"entry_plan_usd_month":null,"hipaa":null,"soc2":null,"gdpr_eu":null,"webrtc":false,"websocket":null,"sip":null,"open_source":false,"self_hostable":true,"price_per_min":null,"custom_avatar":null,"photo_avatar":null,"byo_llm":null,"byo_tts":null,"interruptions":null,"resolution_p":null,"fps":null,"cpu_ok":false,"_notes":"Self-hosted blueprint over RTSP on your own GPUs; production needs NVIDIA AI Enterprise licence. Free preview API credits (up to 5,000) for trials only. Chunk-based, not frame-level.","_added":[]}}]}