Text-to-speech, streaming

Voices that start speaking in a few hundred milliseconds, ideally while the LLM is still writing.

Columns 11/32
Basics
Cost
Features
Performance
Limits
Connects by
Compliance
Openness
Warnings
Compare 0
Must have:

Showing 26 of 26. Click any column header to sort.

xAI Grok TTSxAI
Beta$0.0038$0.013----Yes--1
SpeechifySpeechify
GA$0.0054$0.009$5.26Yes-YesNo5660
Inworld TTSInworld AI
GA$0.0063$0.022$25Yes-YesYes100150
Unreal SpeechUnreal Speech
GA$0.0072$0.015$16.33Yes--No30060
Google Cloud TTSGoogle Cloud
GA$0.009$0.03$30Yes30YesYes--2
Murf Falcon / Falcon 2Murf AI
GA$0.009$0.01$10---Yes100351
ElevenLabsElevenLabs
GA$0.0099$0.072$40Yes-YesYes75322
Soniox Text-to-SpeechSoniox
GA$0.012$0.012$13---Yes-600
Azure Speech TTSMicrosoft
GA$0.013$0.02$15Yes500YesYes300-1
Fish Audio API (S2.1 Pro)Fish Audio
GA$0.013$0.041$15Yes-YesYes90831
OpenAI TTSOpenAI
GA$0.013$0.027$15No-YesNo--1
Amazon PollyAmazon Web Services
GA$0.014$0.027$30Yes-NoYes--1
Mistral Voxtral TTSMistral AI
GA$0.014$0.014$16--Yes-7091
Groq TTS (Canopy Labs Orpheus)Groq
GA$0.02$0.036$22Yes6-No-22
Deepgram Aura-2Deepgram
GA$0.024$0.027$30Yes--Yes20071
Rime TTS (Coda, Mist v3)Rime Labs
GA$0.027$0.045$50Yes--Yes20081
Gradium TTSGradium
Beta$0.032$0.052$57.8Yes--Yes20051
Cartesia SonicCartesia
GA$0.034$0.059$50Yes-YesYes90440
MiniMax Speech (international)MiniMax
GA$0.054$0.09$60--YesYes-401
Neuphonic APINeuphonic
GA-----Yes-25370
Resemble AI TTS (Chatterbox models)Resemble AI
GA-----YesNo20012
Sarvam Bulbul v3Sarvam AI
Beta---Yes--Yes--0
Smallest.ai Waves Lightning v3.1Smallest AI
GA---Yes217YesYes200121
Hume Octave TTSHume AI
Deprecated-----YesYes100111
LMNTLMNT
Shut down---------1
PlayHT (Play.ai)PlayHT (team acqui-hired by Meta)
Shut down---------1

A dash means the vendor does not say. Latency is the vendor's own claim, not our measurement. Per-minute figures are estimates from list prices; each API page explains the basis, because vendors bill by tokens, characters, connection time or flat minutes.

Category warnings

These apply to most APIs in this category.

Vendors are disappearing in 2025-2026

PlayHT's API went offline in July 2025, LMNT's site now says it has shut down, Hume ends its TTS and EVI APIs on 2026-11-13, Groq retired PlayAI TTS, and Rime retired Arcana. Put TTS behind your own interface, keep the source audio for any cloned voice, and keep a second provider tested.

Units differ: characters, credits, UTF-8 bytes, tokens

ElevenLabs and Cartesia bill credits per character, Fish Audio per UTF-8 byte (non-Latin scripts cost ~3x), OpenAI, Gemini-TTS and Soniox per audio token (25 tokens/s for Gemini), Inworld counts UTF-16 code units. Convert everything to cost per minute of your real content before comparing.

Promotional and preview prices expire

ElevenLabs v4/v4 Turbo launch prices end Oct 12 2026 (roughly 3.6x higher afterwards), Gemini 3.8 Flash TTS prices double on Jan 1 2027, Fish Audio's free s2.1-pro-free is reported to end Nov 30 2026, Unreal Speech Basic is $4.99 only for 6 months. Budget on list prices.

Whitespace, markup and SSML are billed

Google counts spaces, newlines and SSML tags (except <mark>); Cartesia charges 1 credit per break tag. Strip markdown, emoji and extra whitespace from LLM output before sending it to TTS.

Vendor latency claims are not your latency

Most headline numbers are model or server-side TTFB excluding network (e.g. Inworld 20 ms P90 server-side, Murf 55 ms model). Independent medians including network are typically 100-450 ms (Vapi Humanness Index, Coval). Measure time-to-first-audio from your own region and also total turn latency.

Chunk LLM output at sentence boundaries

Even with WebSocket input streaming, most engines buffer until punctuation or a character threshold (ElevenLabs chunk_length_schedule starts at 120 chars, MiniMax waits for sentence-final punctuation, ElevenLabs Text to Dialogue waits for ~40 chars and 8 words). Send text at clause/sentence ends and send an explicit flush at end of turn; do not flush every token (Deepgram allows only 20 flushes/min).

HTTP-only providers need your own sentence splitter

OpenAI, Speechify, Unreal Speech, Groq (200-char cap) and Resemble's socket accept a whole text per request. For agents, split replies into sentences and pipeline requests; watch per-request caps and rate limits.

Concurrency limits bite before volume does

Entry plans allow very few simultaneous generations: Smallest.ai 1 per account, Murf 2 outside US-East, Cartesia Free 2 / Pro 3, Fish Audio 5 until $100 spent, Inworld On-Demand 5. Size plans by peak simultaneous speakers.

Voice cloning: consent and disclosure

Clone only voices you have written permission to use. OpenAI requires a recorded consent statement for custom voices and disclosure that the voice is AI-generated; open models (Chatterbox, Dia, Nari Labs) prohibit impersonation. Disclosure duties for synthetic audio also exist in some jurisdictions (for example the EU AI Act transparency rules).

Ask for telephony formats natively

For phone calls request mu-law or A-law 8 kHz from the provider (Cartesia, Deepgram, ElevenLabs, Inworld, Rime, Google, MiniMax support it) instead of resampling. Avoid WAV headers inside streams: they click (Deepgram recommends container=none; Inworld LINEAR16 repeats a header per chunk).

Browsers cannot set WebSocket auth headers

Rime, Deepgram, Inworld, Fish and others expect a header. Proxy through your backend or use the vendor's short-lived tokens (Cartesia access_token, ElevenLabs single_use_token). Never ship a long-lived key to the client; Murf puts the key in the URL, so keep that server-side.

Logging and retention defaults

ElevenLabs enable_logging defaults to true; Fish's free model may retain requests for training. Use zero-retention options (Inworld, Smallest enterprise, Rime retention controls) for sensitive text.