ElevenLabs TTS API
The best-known voice API: Flash v2.5 for low-latency agents over a text-input WebSocket, plus the new Eleven v4 / v4 Turbo (launched 2026-09-28) that stream through a separate Text to Dialogue WebSocket.
Overview
Best for: General-purpose voice agents and content with the largest voice library; Flash v2.5 for latency, v4 for quality.
At a glance
Figures are for Flash v2.5, the model for the streaming WebSocket (~75 ms model latency, 32 languages, $0.04/1K). v4 covers 90+ languages ($0.08/1K list, promo $0.022 until Oct 12 2026). Concurrency 6 is Flash on the Starter plan (Pro 20). Timestamps are character-level alignment. SSML parsing optional on the socket.
Text (SSML parsing optional via enable_ssml_parsing on the WebSocket)
MP3 by default; output_format values follow codec_samplerate_bitrate, e.g. mp3_44100_128, pcm_16000/22050/24000/44100, ulaw_8000, alaw_8000, opus_48000_* (format list from third-party mirrors of the API reference; some higher-quality formats are tier-gated)
Flash v2.5: 32; v3: 70+; v4: 90+
Large shared voice library; instant and professional voice cloning: yes
Vendor claims: Flash v2.5 ~75 ms model latency, v4 Turbo ~100 ms median inference, v3 conversational ~280 ms. All exclude network and application latency.
Global plus residency hosts: api.us.elevenlabs.io, api.eu.residency.elevenlabs.io, api.in.residency.elevenlabs.io, api.sg.residency.elevenlabs.io
Data-residency endpoints for EU, India and Singapore exist. Certifications not re-verified in this pass.
Features
- input streaming (WebSocket)
- multi-context WebSocket for barge-in
- character alignment / timestamps (alignment, sync_alignment)
- voice cloning
- voice design
- pronunciation dictionaries
- data residency endpoints (US, EU, India, Singapore)
Pricing
| What | Price | Unit |
|---|---|---|
| Flash / Turbo | $0.04 | per 1K characters |
| Eleven v4 Turbo | $0.011 | per 1K characters |
| Eleven v4 | $0.022 | per 1K characters |
| Eleven v3 | $0.08 | per 1K characters |
| Eleven v3 Conversational | $0.04 | per 1K characters |
| Multilingual v2 | $0.08 | per 1K characters |
| Starter plan | $6/month ($1 first month) | monthly |
| Creator plan | $22/month | monthly |
| Pro plan | $99/month | monthly |
| Scale plan | $299/month | monthly |
| Business plan | $990/month | monthly |
900 chars/min. Low = v4 Turbo promo rate ($0.011/1K, ends Oct 12 2026; $0.036/min after). Flash v2.5 = $0.036/min. High = v3 or Multilingual v2 at $0.08/1K.
Free tier: Free / pay-as-you-go: 10,000-20,000 characters depending on model. Commercial-use terms on the free tier not re-verified; check the plan terms.
Source: elevenlabs.io
Setup
- Create an API key in the ElevenLabs dashboard.
- Pick a voice_id from the voice library.
- For LLM token streaming with Flash: open wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input?model_id=eleven_flash_v2_5 with the xi-api-key header.
- Send {"text":" "} first, then text chunks ending in a space, then {"text":""} to finish.
- For v4 Turbo use the Text to Dialogue WebSocket instead (different message format: register voices in the first message).
Endpoint
wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input
Authentication
xi-api-key header (or authorization bearer / single_use_token query param for clients)
Quick start javascript
// npm i ws - streams LLM-style text chunks in, saves raw PCM out
import WebSocket from "ws";
import fs from "fs";
const voiceId = "YOUR_VOICE_ID";
const url = `wss://api.elevenlabs.io/v1/text-to-speech/${voiceId}/stream-input?model_id=eleven_flash_v2_5&output_format=pcm_24000`;
const ws = new WebSocket(url, { headers: { "xi-api-key": process.env.ELEVENLABS_API_KEY } });
const out = fs.createWriteStream("out_24k_s16le.pcm");
ws.on("open", () => {
ws.send(JSON.stringify({ text: " " })); // init message
for (const t of ["Hello there. ", "This text arrives ", "in pieces, like LLM tokens. "]) {
ws.send(JSON.stringify({ text: t }));
}
ws.send(JSON.stringify({ text: "" })); // end of input
});
ws.on("message", (raw) => {
const msg = JSON.parse(raw.toString());
if (msg.audio) out.write(Buffer.from(msg.audio, "base64"));
if (msg.isFinal) ws.close();
});
ws.on("close", () => out.end());
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
v3 and v4 do not work on the classic TTS WebSocket
The /text-to-speech/{voice_id}/stream-input socket rejects eleven_v3 and eleven_v4 models. Use eleven_flash_v2_5 there, or switch to the Text to Dialogue WebSocket (wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input) for eleven_v4_turbo, which has a different message format.
v4 launch prices are promotional
The API pricing page shows v4 at $0.022/1K and v4 Turbo at $0.011/1K only until Oct 12 (2026); list prices are $0.08 and $0.04/1K. Budget on the list price.
Default model on the socket is not Flash
If you omit model_id the stream-input socket uses eleven_multilingual_v2 (higher latency, double the price of Flash). Always set model_id explicitly.
Buffering adds latency with small chunks
The server buffers text using chunk_length_schedule (default [120,160,250,290] chars). For conversational agents send flush:true at the end of each turn or tune the schedule, otherwise the first audio waits for 120 characters.
Turbo models are deprecated
eleven_turbo_v2_5 and eleven_turbo_v2 are marked deprecated; migrate to Flash.
Request logging is on by default
enable_logging defaults to true. Zero-retention mode (enable_logging=false) is an enterprise feature per earlier docs; confirm before sending sensitive text.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Chars per request: Flash v2.5 40,000; Multilingual v2 and v4 10,000; v3 5,000 (models page)
- WebSocket inactivity_timeout default 20 s, max 180 s
- Third-party reports (Vapi support): max 5 simultaneous contexts per multi-context WebSocket; not confirmed on official pages
- Plan concurrency limits are not shown on the API pricing page; third-party lists (Free 2 ... Business 15) are unofficial
- Text to Dialogue socket waits for ~40 characters and 8 words before emitting audio unless you flush
Models and products
| Name | Status |
|---|---|
| eleven_v4 | GA (flagship, launched 2026-09-28) |
| eleven_v4_turbo | GA (launched 2026-09-28) |
| eleven_flash_v2_5 | GA |
| eleven_flash_v2 | GA |
| eleven_v3 / eleven_v3_conversational | Previous generation |
| eleven_multilingual_v2 | Previous generation |
| eleven_turbo_v2_5 / eleven_turbo_v2 | Deprecated |
Docs and sources
Docs
Sources used
- elevenlabs.io/pricing/api
- elevenlabs.io/docs/overview/models
- elevenlabs.io/docs/changelog/2026/9/28
- elevenlabs.io/docs/api-reference/text-to-speech/v-1-text-to-speech-voice-id-str...
- elevenlabs.io/docs/eleven-api/guides/how-to/websockets/tts-vs-ttd-websockets
Per-plan concurrency limits; exact output_format list (taken from third-party mirrors); 5-contexts-per-socket limit (third-party); free-tier commercial terms.