GA ElevenLabs

ElevenAgents (formerly Conversational AI / Agents Platform)

A hosted voice-agent platform: ElevenLabs speech recognition, a proprietary turn-taking model, your choice of LLM and ElevenLabs voices, with telephony, RAG and tools. Best for teams that want premium voices and a full no-code/low-code agent builder.

Est. per minute$0.08 - 0.13
2 high-severity warnings

Overview

Best for: Customer-facing agents where voice quality and a managed builder (RAG, tools, telephony, evals) matter more than raw cost.

At a glance

Flat $/min$0.08
Free tierYes
Native S2SNo
ToolsYes
Own LLMYes
Voices5,000
CloningYes
Languages31
Concurrency6
WebSocketYes
Phone / SIPYes
HIPAAYes
EU dataYes
Open weightsNo

Concurrency is the Starter plan (Free 4, Business 40, burst to 3x at $0.16/min). $0.08/min excludes LLM tokens. Free plan 15 min/month. BAA on Enterprise only. Docs also mention 70+ languages for the voice layer.

Audio in

pcm_8000/16000/22050/24000/44100/48000 or ulaw_8000 (negotiated per agent)

Audio out

Same enum as input

Languages

Docs give both '31 languages' and '70+ languages' for the voice layer (inconsistent).

Voices

5k+ voices in the ElevenLabs library, plus voice clones.

Latency

Described as low-latency; no figure on the overview page.

Regions

Default global, plus wss://api.us.elevenlabs.io, EU (api.eu.residency.elevenlabs.io), India (api.in.residency.elevenlabs.io) and Singapore (api.sg.residency.elevenlabs.io) residency hosts.

Compliance

Enterprise: DPA/SLA, BAA for HIPAA, custom SSO; regional data residency hosts (EU, India, Singapore). SOC 2/GDPR not re-verified this session.

Features

  • tool calling (server and client tools)
  • MCP tools with approval
  • RAG knowledge base
  • proprietary turn-taking model
  • interruptions and timeouts config
  • workflow builder
  • multimodal text+voice
  • burst concurrency
  • data residency endpoints

Pricing

WhatPriceUnitNotes
Agent call minutes (all plans)$0.08per minuteDown from $0.10 on 2026-05-07. Billed on connection time; silence over 10 s billed at 5% of the rate.
Burst calls above concurrency limit$0.16per minuteUp to 3x the plan concurrency.
Text messages$0.003per message
LLMprovider list priceper 1M tokensPassed through with no markup; select Gemini and Claude models billed at Vertex regional rates (+10% for US/EU).
Plans (included minutes / concurrency)Free 15/4, Starter $6 75/6, Creator $22 275/10, Pro $99 1,238/20, Scale $299 3,738/30, Business $990 12,375/40per monthEnterprise custom.
Telephonyat costTwilio or SIP carrier bills you directly.
How the per-minute estimate was worked out

$0.08 platform rate plus LLM pass-through; a third-party guide puts LLM at about $0.0005/min (GPT-5 Nano) to $0.045/min (GPT-5.5). Telephony and burst pricing extra.

Audio token rate

Not token-billed (LLM portion is token-billed).

Free tier: 15 call minutes per month, 4 concurrent calls.

Source: elevenlabs.io

Setup

  1. Create an ElevenLabs account and build an agent in the dashboard (prompt, LLM, voice, tools, audio formats).
  2. For public agents connect with ?agent_id; for private agents generate a signed URL server-side.
  3. Stream base64 user_audio_chunk messages, play 'audio' events, answer 'ping' with 'pong'.
  4. For phone calls attach a Twilio number or SIP trunk in the dashboard.

Endpoint

wss://api.elevenlabs.io/v1/convai/conversation?agent_id=<agent_id>

Authentication

agent_id query parameter for public agents; signed URL / token for private agents (xi-api-key for REST)

Quick start javascript

// npm i ws. Create the agent in the dashboard first; set its audio formats to pcm_16000.
// Public agents: agent_id in the URL. Private agents: mint a signed URL on your server.
import WebSocket from "ws";

const ws = new WebSocket(
  `wss://api.elevenlabs.io/v1/convai/conversation?agent_id=${process.env.ELEVEN_AGENT_ID}`
  // EU residency: wss://api.eu.residency.elevenlabs.io/...
);

ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  switch (ev.type) {
    case "conversation_initiation_metadata":
      console.log(ev.conversation_initiation_metadata_event); // negotiated audio formats
      break;
    case "audio":
      play(Buffer.from(ev.audio_event.audio_base_64, "base64"));
      break;
    case "ping":
      ws.send(JSON.stringify({ type: "pong", event_id: ev.ping_event.event_id }));
      break;
  }
});

// 16 kHz mono 16-bit PCM from your mic
export function sendPcm(buf) {
  ws.send(JSON.stringify({ user_audio_chunk: buf.toString("base64") }));
}
function play(pcm) {}

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

LLM is not in the $0.08

LLM tokens are billed on top at provider rates (Gemini/Claude via Vertex regional pricing are 10% higher for US/EU). Large models can add roughly half the platform rate again.

Connection time, not talk time

Billing counts the whole connection; a tab left open or a call never hung up keeps billing (silence over 10 s at 5%). Set max conversation duration and inactivity timeouts.

Burst pricing doubles cost

Calls above plan concurrency are allowed up to 3x but charged $0.16/min. Spiky outbound campaigns can silently double the bill.

Residency limits model choice

With EU data residency some LLMs are unavailable; check model availability per region before committing.

Old rate cards everywhere

Prices fell in May 2026 ($0.10 to $0.08) and included minutes differ across third-party guides; use elevenlabs.io/pricing/agents.

Telephony billed separately

Twilio/SIP minutes are paid to the carrier; ElevenLabs adds no telephony fee.

HIPAA needs Enterprise

BAAs are listed under Enterprise only.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Concurrency 4 (Free) to 40 (Business) concurrent calls; burst to 3x at double price
  • Some LLMs are unavailable when EU data residency is enabled
  • Call time measured on connection duration, not speech time

Models and products

NameStatusNotes
Platform (STT + turn-taking + LLM + ElevenLabs TTS)GACascaded pipeline, not a single speech-to-speech model.
LLM choicesGAOpenAI (GPT-4o mini to GPT-6.1 family), Anthropic (Haiku 4.5 to Opus 5.5), Google Gemini 2.5 to 3.8 Flash, plus ElevenLabs-hosted open models (GLM 5.2, Qwen3.6-35B-A3B, Qwen3.5-397B-A17B, DeepSeek Flash 4.1) or a custom LLM endpoint.

Docs and sources

Docs

Sources used

Not fully verified

Per-minute LLM cost examples come from a third-party guide; WebRTC transport availability not confirmed on the pages read; supported language count is inconsistent in the docs.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?