GA ElevenLabs

ElevenLabs TTS API

The best-known voice API: Flash v2.5 for low-latency agents over a text-input WebSocket, plus the new Eleven v4 / v4 Turbo (launched 2026-09-28) that stream through a separate Text to Dialogue WebSocket.

Est. per minute$0.0099 - 0.072
2 high-severity warnings

Overview

Best for: General-purpose voice agents and content with the largest voice library; Flash v2.5 for latency, v4 for quality.

At a glance

$/1M chars$40
Free tierYes
CloningYes
Text stream inYes
TimestampsYes
SSMLYes
8 kHz phoneYes
Latency ms75
Languages32
Concurrency6
WebRTCNo
WebSocketYes
gRPCNo
EU dataYes
Self-hostNo
Open weightsNo

Figures are for Flash v2.5, the model for the streaming WebSocket (~75 ms model latency, 32 languages, $0.04/1K). v4 covers 90+ languages ($0.08/1K list, promo $0.022 until Oct 12 2026). Concurrency 6 is Flash on the Starter plan (Pro 20). Timestamps are character-level alignment. SSML parsing optional on the socket.

Audio in

Text (SSML parsing optional via enable_ssml_parsing on the WebSocket)

Audio out

MP3 by default; output_format values follow codec_samplerate_bitrate, e.g. mp3_44100_128, pcm_16000/22050/24000/44100, ulaw_8000, alaw_8000, opus_48000_* (format list from third-party mirrors of the API reference; some higher-quality formats are tier-gated)

Languages

Flash v2.5: 32; v3: 70+; v4: 90+

Voices

Large shared voice library; instant and professional voice cloning: yes

Latency

Vendor claims: Flash v2.5 ~75 ms model latency, v4 Turbo ~100 ms median inference, v3 conversational ~280 ms. All exclude network and application latency.

Regions

Global plus residency hosts: api.us.elevenlabs.io, api.eu.residency.elevenlabs.io, api.in.residency.elevenlabs.io, api.sg.residency.elevenlabs.io

Compliance

Data-residency endpoints for EU, India and Singapore exist. Certifications not re-verified in this pass.

Features

  • input streaming (WebSocket)
  • multi-context WebSocket for barge-in
  • character alignment / timestamps (alignment, sync_alignment)
  • voice cloning
  • voice design
  • pronunciation dictionaries
  • data residency endpoints (US, EU, India, Singapore)

Pricing

WhatPriceUnitNotes
Flash / Turbo$0.04per 1K charactersAPI pricing page; ~$0.04/min per vendor
Eleven v4 Turbo$0.011per 1K charactersPromotional, discounted from $0.04 until Oct 12 (2026)
Eleven v4$0.022per 1K charactersPromotional, discounted from $0.08 until Oct 12 (2026)
Eleven v3$0.08per 1K characters
Eleven v3 Conversational$0.04per 1K characters
Multilingual v2$0.08per 1K characters
Starter plan$6/month ($1 first month)monthly150,000 Flash chars or 75,000 v3 chars
Creator plan$22/monthmonthly550,000 Flash chars
Pro plan$99/monthmonthly2,475,000 Flash chars
Scale plan$299/monthmonthly7,475,000 Flash chars
Business plan$990/monthmonthly24,750,000 Flash chars
How the per-minute estimate was worked out

900 chars/min. Low = v4 Turbo promo rate ($0.011/1K, ends Oct 12 2026; $0.036/min after). Flash v2.5 = $0.036/min. High = v3 or Multilingual v2 at $0.08/1K.

Free tier: Free / pay-as-you-go: 10,000-20,000 characters depending on model. Commercial-use terms on the free tier not re-verified; check the plan terms.

Source: elevenlabs.io

Setup

  1. Create an API key in the ElevenLabs dashboard.
  2. Pick a voice_id from the voice library.
  3. For LLM token streaming with Flash: open wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input?model_id=eleven_flash_v2_5 with the xi-api-key header.
  4. Send {"text":" "} first, then text chunks ending in a space, then {"text":""} to finish.
  5. For v4 Turbo use the Text to Dialogue WebSocket instead (different message format: register voices in the first message).

Endpoint

wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input

Authentication

xi-api-key header (or authorization bearer / single_use_token query param for clients)

Quick start javascript

// npm i ws  - streams LLM-style text chunks in, saves raw PCM out
import WebSocket from "ws";
import fs from "fs";

const voiceId = "YOUR_VOICE_ID";
const url = `wss://api.elevenlabs.io/v1/text-to-speech/${voiceId}/stream-input?model_id=eleven_flash_v2_5&output_format=pcm_24000`;
const ws = new WebSocket(url, { headers: { "xi-api-key": process.env.ELEVENLABS_API_KEY } });
const out = fs.createWriteStream("out_24k_s16le.pcm");

ws.on("open", () => {
  ws.send(JSON.stringify({ text: " " }));                 // init message
  for (const t of ["Hello there. ", "This text arrives ", "in pieces, like LLM tokens. "]) {
    ws.send(JSON.stringify({ text: t }));
  }
  ws.send(JSON.stringify({ text: "" }));                  // end of input
});
ws.on("message", (raw) => {
  const msg = JSON.parse(raw.toString());
  if (msg.audio) out.write(Buffer.from(msg.audio, "base64"));
  if (msg.isFinal) ws.close();
});
ws.on("close", () => out.end());

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

v3 and v4 do not work on the classic TTS WebSocket

The /text-to-speech/{voice_id}/stream-input socket rejects eleven_v3 and eleven_v4 models. Use eleven_flash_v2_5 there, or switch to the Text to Dialogue WebSocket (wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input) for eleven_v4_turbo, which has a different message format.

v4 launch prices are promotional

The API pricing page shows v4 at $0.022/1K and v4 Turbo at $0.011/1K only until Oct 12 (2026); list prices are $0.08 and $0.04/1K. Budget on the list price.

Default model on the socket is not Flash

If you omit model_id the stream-input socket uses eleven_multilingual_v2 (higher latency, double the price of Flash). Always set model_id explicitly.

Buffering adds latency with small chunks

The server buffers text using chunk_length_schedule (default [120,160,250,290] chars). For conversational agents send flush:true at the end of each turn or tune the schedule, otherwise the first audio waits for 120 characters.

Turbo models are deprecated

eleven_turbo_v2_5 and eleven_turbo_v2 are marked deprecated; migrate to Flash.

Request logging is on by default

enable_logging defaults to true. Zero-retention mode (enable_logging=false) is an enterprise feature per earlier docs; confirm before sending sensitive text.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • Chars per request: Flash v2.5 40,000; Multilingual v2 and v4 10,000; v3 5,000 (models page)
  • WebSocket inactivity_timeout default 20 s, max 180 s
  • Third-party reports (Vapi support): max 5 simultaneous contexts per multi-context WebSocket; not confirmed on official pages
  • Plan concurrency limits are not shown on the API pricing page; third-party lists (Free 2 ... Business 15) are unofficial
  • Text to Dialogue socket waits for ~40 characters and 8 words before emitting audio unless you flush

Models and products

NameStatusNotes
eleven_v4GA (flagship, launched 2026-09-28)Highest quality, 90+ languages, 10,000 chars/request. Used via the Text to Dialogue API, not the classic TTS WebSocket.
eleven_v4_turboGA (launched 2026-09-28)~100 ms median inference (vendor). For agents via wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input. One voice per connection.
eleven_flash_v2_5GA~75 ms model latency (vendor), 32 languages, 40,000 chars/request. Recommended for the classic stream-input WebSocket.
eleven_flash_v2GAEnglish only, ~75 ms (vendor).
eleven_v3 / eleven_v3_conversationalPrevious generation70+ languages; v3 conversational ~280 ms. Not accepted on the classic TTS WebSocket; use Text to Dialogue.
eleven_multilingual_v2Previous generation29 languages, 10,000 chars/request. Default model_id of the stream-input WebSocket if you omit it.
eleven_turbo_v2_5 / eleven_turbo_v2DeprecatedDocs say use Flash instead; functionally equivalent.

Docs and sources

Docs

Sources used

Not fully verified

Per-plan concurrency limits; exact output_format list (taken from third-party mirrors); 5-contexts-per-socket limit (third-party); free-tier commercial terms.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?