GA ElevenLabs

ElevenLabs Scribe v2 Realtime

WebSocket realtime STT (scribe_v2_realtime) with VAD or manual commits, 90+ languages, keyterms and entity detection, billed from your ElevenLabs plan credits.

Est. per minute$0.0065 - 0.0098

Overview

Best for: Teams already using ElevenLabs TTS/agents who want the same vendor for live STT, including 8 kHz mu-law telephony.

At a glance

$/hour$0.39
Free tierYes
Turn detectYes
KeytermsYes
PartialsYes
8 kHz phoneYes
Latency ms150
Languages90
Concurrency9
WebRTCNo
WebSocketYes
gRPCNo
EU dataYes
Self-hostNo
Open weightsNo

~150 ms vendor claim, conditions not stated. 90+ languages. Concurrency 9 is the Starter plan; Pro 30, Scale 45, Free 6. Free plan includes about 2.5 realtime hours. Diarization not listed for realtime. Keyterms +$0.08/hr. EU, India and Singapore residency hosts.

Audio in

audio_format: pcm_8000, pcm_16000 (default), pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000; audio sent base64 in JSON input_audio_chunk messages.

Audio out

n/a

Languages

90+ languages (vendor); language_code plus secondary_languages hints.

Latency

Vendor claim: ~150 ms (footnoted, conditions not stated).

Regions

Default api.elevenlabs.io; US host and EU, India, Singapore data-residency hosts listed.

Compliance

Data-residency hosts for EU, India, Singapore; enable_logging=false for zero retention. Certifications not re-verified.

Features

  • partial_transcript and committed_transcript
  • commit_strategy: manual, vad, or turn_prediction
  • VAD tuning (threshold, silence secs, min speech/silence ms)
  • word timestamps (include_timestamps)
  • language detection
  • keyterms (up to 50, 20 chars each; paid)
  • entity detection
  • experimental transcript editing (paid)
  • background audio filter
  • enable_logging=false option
  • regional and data-residency hosts

Pricing

WhatPriceUnitNotes
Scribe v2 Realtime$0.39per hourSame rate on every plan; included hours scale with plan (Free 2h30, Starter 15h, Creator 56h, Pro 254h, Scale 767h, Business 2,538h).
Keyterm prompting (realtime)+$0.08per hour
Transcript editing (realtime)+$0.12per hourExperimental.
Scribe v2 batch (reference)$0.22per hour
How the per-minute estimate was worked out

$0.39/hr base to $0.59/hr with keyterms + editing.

Free tier: Free plan includes about 2.5 realtime hours.

Source: elevenlabs.io

Setup

  1. Create an API key.
  2. Server: connect to wss://api.elevenlabs.io/v1/speech-to-text/realtime?model_id=scribe_v2_realtime&audio_format=pcm_16000&commit_strategy=vad with header xi-api-key.
  3. Browser: create a single-use token via the tokens endpoint on your server and pass ?token=.
  4. Send input_audio_chunk JSON messages with base64 audio; read partial_transcript and committed_transcript.

Endpoint

wss://api.elevenlabs.io/v1/speech-to-text/realtime

Authentication

xi-api-key header, or single-use token in the token query parameter for client-side use.

Quick start python

import asyncio, base64, json, os, websockets  # pip install websockets>=14

URL = ("wss://api.elevenlabs.io/v1/speech-to-text/realtime"
       "?model_id=scribe_v2_realtime&audio_format=pcm_16000&commit_strategy=vad")

async def main():
    hdr = {"xi-api-key": os.environ["ELEVENLABS_API_KEY"]}
    async with websockets.connect(URL, additional_headers=hdr) as ws:
        async def send():
            with open("audio_16k_mono.raw", "rb") as f:
                while chunk := f.read(3200):  # 100 ms
                    await ws.send(json.dumps({
                        "message_type": "input_audio_chunk",
                        "audio_base_64": base64.b64encode(chunk).decode(),
                        "commit": False, "sample_rate": 16000}))
                    await asyncio.sleep(0.1)
        asyncio.create_task(send())
        async for msg in ws:
            d = json.loads(msg)
            t = d.get("message_type")
            if t == "partial_transcript":
                print("partial", d.get("text"))
            elif t == "committed_transcript":
                print("FINAL", d.get("text"))
            elif t and t.endswith("error"):
                print(d); break

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No realtime diarization listed

Speaker diarization is documented for batch Scribe v2 only. For multi-party calls, send separate channels as separate sessions.

Logging on by default

enable_logging defaults to true; set it to false if you need no retention (may require an eligible plan).

Base64 JSON overhead

Audio goes as base64 inside JSON (about 33% larger than binary frames), which matters on mobile uplinks.

Plan credits, not pure PAYG

Realtime STT draws from subscription credits; included hours depend on the plan, and overage behaviour follows your plan settings.

Keyterm limits are small

Realtime supports at most 50 keyterms of 20 characters each, and they cost +$0.08/hr.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Realtime concurrency limit by plan is in a separate chart on the models page (not captured).
  • keepalive_interval_ms configurable 500-10000; errors include session_time_limit_exceeded, commit_throttled, queue_overflow, insufficient_audio_activity.

Models and products

NameStatusNotes
scribe_v2_realtimeGA (only realtime model)Batch counterpart is Scribe v2 (has diarization); realtime does not list diarization.

Docs and sources

Docs

Sources used

Not fully verified

Realtime concurrency per plan, whether enable_logging=false is limited to enterprise, exact error message_type names used in the snippet's error check, and the meaning of the ~150 ms footnote.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?