Soniox Real-time STT (stt-rt-v5)
Very cheap token-priced realtime STT over WebSocket with 60+ languages, speaker separation, language ID, endpoint detection and realtime translation included in one model.
Overview
Best for: Cost-sensitive multilingual live transcription and live translation, with diarization and language ID without add-on fees.
At a glance
Token-billed; $0.12/hr is the vendor calculator estimate with diarization and language ID included. 60+ languages. 300-minute stream cap cannot be raised. Concurrency raisable in console.
audio_format 'auto' for containers, or raw PCM: pcm_s8/s16/s24/s32, unsigned u8-u32, float f32/f64 (le/be), with sample_rate and num_channels.
n/a
60+ languages with language identification and mixed-language robustness; realtime translation across 3,600+ language pairs (vendor).
No numeric vendor claim captured; model notes cite 'faster semantic endpointing'.
Not re-verified in this pass.
Not re-verified in this pass.
Features
- non-final and final tokens
- semantic endpoint detection (<end> token, tunable sensitivity and max delay)
- manual finalize (<fin> token)
- speaker separation (diarization) included
- language identification included
- context / custom vocabulary
- realtime translation
- keepalive control message
- temporary API keys for browsers
Pricing
| What | Price | Unit |
|---|---|---|
| Real-time input audio | $2.00 | per 1M audio tokens |
| Real-time input text (context) | $4.00 | per 1M tokens |
| Real-time output text | $4.00 | per 1M tokens |
| Real-time effective rate | ~$0.12 | per hour |
~$0.12/hr vendor estimate = $0.002/min; higher with long context prompts or translation output.
Free tier: Not stated on the pricing page.
Source: soniox.com
Setup
- Create an API key in the Soniox Console.
- Connect to wss://stt-rt.soniox.com/transcribe-websocket with Authorization: Bearer <key> (or a temporary key for browsers).
- Send a JSON start message with model, audio_format, sample_rate, num_channels and options.
- Stream binary audio; rebuild text from tokens (append is_final tokens, redraw non-final ones).
- Send an empty TEXT frame to finish; wait for {"finished": true}.
Endpoint
wss://stt-rt.soniox.com/transcribe-websocket
Authentication
Authorization: Bearer <API_KEY> header (api_key in the start message is deprecated); temporary API keys for client-side use.
Quick start python
import asyncio, json, os, websockets # pip install websockets>=14
URL = "wss://stt-rt.soniox.com/transcribe-websocket"
CFG = {"model": "stt-rt-v5", "audio_format": "pcm_s16le", "sample_rate": 16000,
"num_channels": 1, "enable_endpoint_detection": True}
async def main():
hdr = {"Authorization": f"Bearer {os.environ['SONIOX_API_KEY']}"}
async with websockets.connect(URL, additional_headers=hdr) as ws:
await ws.send(json.dumps(CFG))
async def send():
with open("audio_16k_mono.raw", "rb") as f:
while chunk := f.read(3200): # 100 ms
await ws.send(chunk)
await asyncio.sleep(0.1)
await ws.send("") # empty TEXT frame ends the stream
asyncio.create_task(send())
final = ""
async for msg in ws:
d = json.loads(msg)
if d.get("error_code"):
print(d); break
toks = d.get("tokens", [])
final += "".join(t["text"] for t in toks if t["is_final"])
partial = "".join(t["text"] for t in toks if not t["is_final"])
print(final + " | " + partial)
if d.get("finished"):
break
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Default concurrency is only 10
New accounts get 10 concurrent streams and 100 requests per minute. Raise it in the console well before a launch.
Hard 300-minute stream cap
Each session ends at 300 minutes and Soniox says this cannot be increased; long broadcasts must roll over to a new session.
Empty binary frame does not end the stream
Only an empty TEXT frame finishes the session; an empty binary frame is treated as an empty audio chunk and the session stays open.
Token-based bill is an estimate
Pricing is per audio, context and output token; the ~$0.12/hr figure assumes typical speech density. Large context prompts and translation output add text tokens.
Model auto-upgrades via aliases
stt-rt-v4 was silently routed to stt-rt-v5 after 2026-06-30. If you need stable behaviour for evaluation, pin the exact model and watch the changelog.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- 100 requests per minute, 10 concurrent streams by default (raisable in console).
- Each stream capped at 300 minutes; this cap cannot be raised.
- Open-connection limit is a multiple of the concurrent-request limit, including idle sockets.
Models and products
| Name | Status |
|---|---|
| stt-rt-v5 | Active (released 2026-06-16) |
| stt-async-v5 | Active |
Docs and sources
Docs
Sources used
- soniox.com/pricing
- soniox.com/docs/stt/models.mdx
- soniox.com/docs/stt/rt/limits-and-quotas.mdx
- soniox.com/docs/stt/rt/real-time-transcription.mdx
- soniox.com/docs/api-reference/stt/websocket-api.mdx
Free credit, regions and compliance.