AssemblyAI Universal-3.6 Pro Realtime and Universal-Streaming
WebSocket streaming STT (API v3) with built-in turn detection. Universal-3.6 Pro is the default high-accuracy multilingual streaming model; Universal-Streaming English/Multilingual are the cheaper tier.
Overview
Best for: Voice agents and live captioning that want turn detection built in, multilingual code-switching, and US/EU data residency on a simple hourly price.
At a glance
Price and languages are for the default Universal-3.6 Pro model; Universal-Streaming is $0.15/hr (English, or 6 languages). Limit is 100+ new sessions per minute on paid, not a concurrent-stream cap. Diarization +$0.12/hr. Streaming PII pricing unconfirmed. HIPAA and SOC 2 are vendor claims, not re-verified.
Mono 16-bit PCM by default with sample_rate set to the source; mulaw also documented historically; AAC (ADTS, encoding=aac) and Opus (encoding=ogg_opus or opus) accepted. Chunks must be 50-1000 ms and must not arrive faster than real time.
n/a
Universal-3.6 Pro: 32 languages with code-switching; 3.5 Pro: 19; Universal-Streaming Multilingual: 6; English model: English only.
Vendor docs rate the Pro models 'Fastest' and Universal-Streaming 'Fast'; no numeric figure on the model page.
Edge routing across AWS US and EU regions by default; data-zone endpoints keep data in US (streaming.us.assemblyai.com) or EU (streaming.eu.assemblyai.com). Self-hosted streaming offered.
US and EU data-zone streaming endpoints documented. Vendor markets SOC 2 and HIPAA BAA; not re-verified in this pass.
Features
- partial and final turns (Turn messages with end_of_turn)
- turn detection with min_turn_silence / max_turn_silence / vad_threshold and latency modes
- keyterm prompting (up to 100 terms)
- general prompting (beta, Pro)
- streaming diarization and multichannel (paid)
- medical mode (paid)
- Voice Focus noise isolation (paid, Pro)
- PII redaction and profanity filter pages for streaming
- UpdateConfiguration mid-stream
- session-end webhooks
- US/EU data-zone endpoints
Pricing
| What | Price | Unit |
|---|---|---|
| Universal-3.6 Pro Realtime | $0.45 | per hour of session |
| Universal-Streaming English | $0.15 | per hour of session |
| Universal-Streaming Multilingual | $0.15 | per hour of session |
| Streaming speaker diarization | +$0.12 | per hour |
| Medical Mode | +$0.15 | per hour |
| General prompting (beta) | +$0.05 | per hour |
| Voice Focus | +$0.10 | per hour |
$0.15/hr Universal-Streaming = $0.0025/min; $0.45/hr U3.6 Pro + $0.12/hr diarization = $0.0095/min. Session time, so idle open sockets cost the same.
Free tier: $50 free credit at signup, no card required.
Source: assemblyai.com
Setup
- Get an API key from the dashboard.
- Open wss://streaming.assemblyai.com/v3/ws?speech_model=universal-3-6-pro&sample_rate=16000 with header Authorization: <API_KEY>.
- Stream 50-1000 ms binary chunks at real-time pace; read Begin, Turn and Termination messages.
- Send {"type":"Terminate"} when the call ends, otherwise you keep paying.
- For browsers, mint a temporary token server-side (expires_in_seconds, optional max_session_duration_seconds) and pass it as token= in the URL.
Endpoint
wss://streaming.assemblyai.com/v3/ws (US: wss://streaming.us.assemblyai.com/v3/ws, EU: wss://streaming.eu.assemblyai.com/v3/ws)
Authentication
Authorization header with the raw API key; browser clients use a temporary streaming token generated server-side.
Quick start python
import asyncio, json, os, websockets # pip install websockets>=14
from urllib.parse import urlencode
params = {"speech_model": "universal-3-6-pro", "sample_rate": 16000}
URL = "wss://streaming.assemblyai.com/v3/ws?" + urlencode(params)
async def main():
hdr = {"Authorization": os.environ["ASSEMBLYAI_API_KEY"]}
async with websockets.connect(URL, additional_headers=hdr) as ws:
async def send():
with open("audio_16k_mono.raw", "rb") as f: # 16-bit PCM mono
while chunk := f.read(3200): # 100 ms (must be 50-1000 ms)
await ws.send(chunk)
await asyncio.sleep(0.1) # never faster than real time
await ws.send(json.dumps({"type": "Terminate"})) # stops billing
asyncio.create_task(send())
async for msg in ws:
d = json.loads(msg)
if d["type"] == "Turn":
print("FINAL" if d.get("end_of_turn") else "partial", d["transcript"])
elif d["type"] == "Termination":
break
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Billed for socket open time
Streaming is billed per session duration, the time the WebSocket is open, not the audio you send. Idle time counts and add-ons like diarization apply to the whole session. Always send Terminate when the call ends.
Forgotten sessions bill for 3 hours
Sessions that are never terminated auto-close after 3 hours and are billed for the full duration; they also keep counting against your new-session rate limit.
Strict chunk size and pacing
Chunks under 50 ms or over 1000 ms close the session with 3007, and so does sending audio faster than real time. Streaming a file needs a sleep between chunks.
Model IDs changed in 2026
Older u3-rt-pro / u3-pro streaming IDs are no longer in the API reference, and the default model is now universal-3-6-pro (more expensive than Universal-Streaming). Pin speech_model explicitly so cost does not change under you.
Guardrail add-ons not confirmed for streaming pricing
The pricing page lists PII redaction/profanity/content moderation prices only for batch models, although streaming docs have PII and profanity pages. Confirm cost before enabling.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- Rate limit is on NEW sessions per minute, not total open sessions: Free 5/min, paid 100+/min, auto-scales +10% when you use 70%+ of it.
- Max session 3 hours (error 3008); sessions not terminated auto-close after 3 hours and are billed for the full time.
- Audio chunk must be 50-1000 ms (error 3007); sending faster than real time is rejected (3007).
- Too many sessions gives close code 1008/3009.
Models and products
| Name | Status |
|---|---|
| universal-3-6-pro | GA (default) |
| universal-3-5-pro | Listed in docs |
| universal-streaming-english | GA |
| universal-streaming-multilingual | GA |
| u3-rt-pro / u3-pro | Removed from current API reference |
Docs and sources
Docs
Sources used
- assemblyai.com/pricing
- assemblyai.com/docs/streaming/universal-streaming
- assemblyai.com/docs/streaming/endpoints-and-data-zones.md
- assemblyai.com/docs/streaming/rate-limits.md
- assemblyai.com/docs/streaming/common-session-errors-and-closures.md
- assemblyai.com/docs/streaming/turn-detection.md
Numeric latency, compliance certifications, and exact temporary-token endpoint path (docs show GET generate-streaming-token reference) not re-checked.