GA AssemblyAI

AssemblyAI Universal-3.6 Pro Realtime and Universal-Streaming

WebSocket streaming STT (API v3) with built-in turn detection. Universal-3.6 Pro is the default high-accuracy multilingual streaming model; Universal-Streaming English/Multilingual are the cheaper tier.

Est. per minute$0.0025 - 0.0095
2 high-severity warnings

Overview

Best for: Voice agents and live captioning that want turn detection built in, multilingual code-switching, and US/EU data residency on a simple hourly price.

At a glance

$/hour$0.45
Bills silenceYes
Free tierYes
Free credit $$50
Live speakersYes
Turn detectYes
KeytermsYes
PII redactYes
PartialsYes
Mixed langsYes
Languages32
Max session min180
WebRTCNo
WebSocketYes
gRPCNo
HIPAAYes
SOC 2Yes
EU dataYes
Self-hostYes
Open weightsNo

Price and languages are for the default Universal-3.6 Pro model; Universal-Streaming is $0.15/hr (English, or 6 languages). Limit is 100+ new sessions per minute on paid, not a concurrent-stream cap. Diarization +$0.12/hr. Streaming PII pricing unconfirmed. HIPAA and SOC 2 are vendor claims, not re-verified.

Audio in

Mono 16-bit PCM by default with sample_rate set to the source; mulaw also documented historically; AAC (ADTS, encoding=aac) and Opus (encoding=ogg_opus or opus) accepted. Chunks must be 50-1000 ms and must not arrive faster than real time.

Audio out

n/a

Languages

Universal-3.6 Pro: 32 languages with code-switching; 3.5 Pro: 19; Universal-Streaming Multilingual: 6; English model: English only.

Latency

Vendor docs rate the Pro models 'Fastest' and Universal-Streaming 'Fast'; no numeric figure on the model page.

Regions

Edge routing across AWS US and EU regions by default; data-zone endpoints keep data in US (streaming.us.assemblyai.com) or EU (streaming.eu.assemblyai.com). Self-hosted streaming offered.

Compliance

US and EU data-zone streaming endpoints documented. Vendor markets SOC 2 and HIPAA BAA; not re-verified in this pass.

Features

  • partial and final turns (Turn messages with end_of_turn)
  • turn detection with min_turn_silence / max_turn_silence / vad_threshold and latency modes
  • keyterm prompting (up to 100 terms)
  • general prompting (beta, Pro)
  • streaming diarization and multichannel (paid)
  • medical mode (paid)
  • Voice Focus noise isolation (paid, Pro)
  • PII redaction and profanity filter pages for streaming
  • UpdateConfiguration mid-stream
  • session-end webhooks
  • US/EU data-zone endpoints

Pricing

WhatPriceUnitNotes
Universal-3.6 Pro Realtime$0.45per hour of sessionBilled on WebSocket open time, not audio sent.
Universal-Streaming English$0.15per hour of sessionKeyterms +$0.04/hr on this model.
Universal-Streaming Multilingual$0.15per hour of session
Streaming speaker diarization+$0.12per hourApplies to the whole session.
Medical Mode+$0.15per hour
General prompting (beta)+$0.05per hourUniversal-3.6 Pro only.
Voice Focus+$0.10per hourUniversal-3.6 Pro only.
How the per-minute estimate was worked out

$0.15/hr Universal-Streaming = $0.0025/min; $0.45/hr U3.6 Pro + $0.12/hr diarization = $0.0095/min. Session time, so idle open sockets cost the same.

Free tier: $50 free credit at signup, no card required.

Source: assemblyai.com

Setup

  1. Get an API key from the dashboard.
  2. Open wss://streaming.assemblyai.com/v3/ws?speech_model=universal-3-6-pro&sample_rate=16000 with header Authorization: <API_KEY>.
  3. Stream 50-1000 ms binary chunks at real-time pace; read Begin, Turn and Termination messages.
  4. Send {"type":"Terminate"} when the call ends, otherwise you keep paying.
  5. For browsers, mint a temporary token server-side (expires_in_seconds, optional max_session_duration_seconds) and pass it as token= in the URL.

Endpoint

wss://streaming.assemblyai.com/v3/ws (US: wss://streaming.us.assemblyai.com/v3/ws, EU: wss://streaming.eu.assemblyai.com/v3/ws)

Authentication

Authorization header with the raw API key; browser clients use a temporary streaming token generated server-side.

Quick start python

import asyncio, json, os, websockets  # pip install websockets>=14
from urllib.parse import urlencode

params = {"speech_model": "universal-3-6-pro", "sample_rate": 16000}
URL = "wss://streaming.assemblyai.com/v3/ws?" + urlencode(params)

async def main():
    hdr = {"Authorization": os.environ["ASSEMBLYAI_API_KEY"]}
    async with websockets.connect(URL, additional_headers=hdr) as ws:
        async def send():
            with open("audio_16k_mono.raw", "rb") as f:  # 16-bit PCM mono
                while chunk := f.read(3200):  # 100 ms (must be 50-1000 ms)
                    await ws.send(chunk)
                    await asyncio.sleep(0.1)  # never faster than real time
            await ws.send(json.dumps({"type": "Terminate"}))  # stops billing
        asyncio.create_task(send())
        async for msg in ws:
            d = json.loads(msg)
            if d["type"] == "Turn":
                print("FINAL" if d.get("end_of_turn") else "partial", d["transcript"])
            elif d["type"] == "Termination":
                break

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Billed for socket open time

Streaming is billed per session duration, the time the WebSocket is open, not the audio you send. Idle time counts and add-ons like diarization apply to the whole session. Always send Terminate when the call ends.

Forgotten sessions bill for 3 hours

Sessions that are never terminated auto-close after 3 hours and are billed for the full duration; they also keep counting against your new-session rate limit.

Strict chunk size and pacing

Chunks under 50 ms or over 1000 ms close the session with 3007, and so does sending audio faster than real time. Streaming a file needs a sleep between chunks.

Model IDs changed in 2026

Older u3-rt-pro / u3-pro streaming IDs are no longer in the API reference, and the default model is now universal-3-6-pro (more expensive than Universal-Streaming). Pin speech_model explicitly so cost does not change under you.

Guardrail add-ons not confirmed for streaming pricing

The pricing page lists PII redaction/profanity/content moderation prices only for batch models, although streaming docs have PII and profanity pages. Confirm cost before enabling.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Rate limit is on NEW sessions per minute, not total open sessions: Free 5/min, paid 100+/min, auto-scales +10% when you use 70%+ of it.
  • Max session 3 hours (error 3008); sessions not terminated auto-close after 3 hours and are billed for the full time.
  • Audio chunk must be 50-1000 ms (error 3007); sending faster than real time is rejected (3007).
  • Too many sessions gives close code 1008/3009.

Models and products

NameStatusNotes
universal-3-6-proGA (default)Default if speech_model is omitted. 32 languages with native code-switching. Keyterms included, prompting (beta) and Voice Focus as paid add-ons.
universal-3-5-proListed in docs19 languages.
universal-streaming-englishGAEnglish only, cheapest tier.
universal-streaming-multilingualGAEnglish, Spanish, German, French, Portuguese, Italian.
u3-rt-pro / u3-proRemoved from current API referenceOlder Universal-3 streaming IDs per pricing page note; migrate to universal-3-6-pro.

Docs and sources

Docs

Sources used

Not fully verified

Numeric latency, compliance certifications, and exact temporary-token endpoint path (docs show GET generate-streaming-token reference) not re-checked.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?