ElevenLabs Scribe v2 Realtime
WebSocket realtime STT (scribe_v2_realtime) with VAD or manual commits, 90+ languages, keyterms and entity detection, billed from your ElevenLabs plan credits.
Overview
Best for: Teams already using ElevenLabs TTS/agents who want the same vendor for live STT, including 8 kHz mu-law telephony.
At a glance
~150 ms vendor claim, conditions not stated. 90+ languages. Concurrency 9 is the Starter plan; Pro 30, Scale 45, Free 6. Free plan includes about 2.5 realtime hours. Diarization not listed for realtime. Keyterms +$0.08/hr. EU, India and Singapore residency hosts.
audio_format: pcm_8000, pcm_16000 (default), pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000; audio sent base64 in JSON input_audio_chunk messages.
n/a
90+ languages (vendor); language_code plus secondary_languages hints.
Vendor claim: ~150 ms (footnoted, conditions not stated).
Default api.elevenlabs.io; US host and EU, India, Singapore data-residency hosts listed.
Data-residency hosts for EU, India, Singapore; enable_logging=false for zero retention. Certifications not re-verified.
Features
- partial_transcript and committed_transcript
- commit_strategy: manual, vad, or turn_prediction
- VAD tuning (threshold, silence secs, min speech/silence ms)
- word timestamps (include_timestamps)
- language detection
- keyterms (up to 50, 20 chars each; paid)
- entity detection
- experimental transcript editing (paid)
- background audio filter
- enable_logging=false option
- regional and data-residency hosts
Pricing
| What | Price | Unit |
|---|---|---|
| Scribe v2 Realtime | $0.39 | per hour |
| Keyterm prompting (realtime) | +$0.08 | per hour |
| Transcript editing (realtime) | +$0.12 | per hour |
| Scribe v2 batch (reference) | $0.22 | per hour |
$0.39/hr base to $0.59/hr with keyterms + editing.
Free tier: Free plan includes about 2.5 realtime hours.
Source: elevenlabs.io
Setup
- Create an API key.
- Server: connect to wss://api.elevenlabs.io/v1/speech-to-text/realtime?model_id=scribe_v2_realtime&audio_format=pcm_16000&commit_strategy=vad with header xi-api-key.
- Browser: create a single-use token via the tokens endpoint on your server and pass ?token=.
- Send input_audio_chunk JSON messages with base64 audio; read partial_transcript and committed_transcript.
Endpoint
wss://api.elevenlabs.io/v1/speech-to-text/realtime
Authentication
xi-api-key header, or single-use token in the token query parameter for client-side use.
Quick start python
import asyncio, base64, json, os, websockets # pip install websockets>=14
URL = ("wss://api.elevenlabs.io/v1/speech-to-text/realtime"
"?model_id=scribe_v2_realtime&audio_format=pcm_16000&commit_strategy=vad")
async def main():
hdr = {"xi-api-key": os.environ["ELEVENLABS_API_KEY"]}
async with websockets.connect(URL, additional_headers=hdr) as ws:
async def send():
with open("audio_16k_mono.raw", "rb") as f:
while chunk := f.read(3200): # 100 ms
await ws.send(json.dumps({
"message_type": "input_audio_chunk",
"audio_base_64": base64.b64encode(chunk).decode(),
"commit": False, "sample_rate": 16000}))
await asyncio.sleep(0.1)
asyncio.create_task(send())
async for msg in ws:
d = json.loads(msg)
t = d.get("message_type")
if t == "partial_transcript":
print("partial", d.get("text"))
elif t == "committed_transcript":
print("FINAL", d.get("text"))
elif t and t.endswith("error"):
print(d); break
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
No realtime diarization listed
Speaker diarization is documented for batch Scribe v2 only. For multi-party calls, send separate channels as separate sessions.
Logging on by default
enable_logging defaults to true; set it to false if you need no retention (may require an eligible plan).
Base64 JSON overhead
Audio goes as base64 inside JSON (about 33% larger than binary frames), which matters on mobile uplinks.
Plan credits, not pure PAYG
Realtime STT draws from subscription credits; included hours depend on the plan, and overage behaviour follows your plan settings.
Keyterm limits are small
Realtime supports at most 50 keyterms of 20 characters each, and they cost +$0.08/hr.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- Realtime concurrency limit by plan is in a separate chart on the models page (not captured).
- keepalive_interval_ms configurable 500-10000; errors include session_time_limit_exceeded, commit_throttled, queue_overflow, insufficient_audio_activity.
Models and products
| Name | Status |
|---|---|
| scribe_v2_realtime | GA (only realtime model) |
Docs and sources
Docs
Sources used
- elevenlabs.io/pricing/api
- elevenlabs.io/docs/capabilities/speech-to-text
- elevenlabs.io/docs/api-reference/speech-to-text/v-1-speech-to-text-realtime
Realtime concurrency per plan, whether enable_logging=false is limited to enterprise, exact error message_type names used in the snippet's error check, and the meaning of the ~150 ms footnote.