GA Soniox

Soniox Real-time STT (stt-rt-v5)

Very cheap token-priced realtime STT over WebSocket with 60+ languages, speaker separation, language ID, endpoint detection and realtime translation included in one model.

Est. per minute$0.002 - 0.0025
1 high-severity warning

Overview

Best for: Cost-sensitive multilingual live transcription and live translation, with diarization and language ID without add-on fees.

At a glance

$/hour$0.12
Live speakersYes
Turn detectYes
KeytermsYes
PartialsYes
Mixed langsYes
Languages60
Max session min300
Concurrency10
WebRTCNo
WebSocketYes
gRPCNo
Self-hostNo
Open weightsNo

Token-billed; $0.12/hr is the vendor calculator estimate with diarization and language ID included. 60+ languages. 300-minute stream cap cannot be raised. Concurrency raisable in console.

Audio in

audio_format 'auto' for containers, or raw PCM: pcm_s8/s16/s24/s32, unsigned u8-u32, float f32/f64 (le/be), with sample_rate and num_channels.

Audio out

n/a

Languages

60+ languages with language identification and mixed-language robustness; realtime translation across 3,600+ language pairs (vendor).

Latency

No numeric vendor claim captured; model notes cite 'faster semantic endpointing'.

Regions

Not re-verified in this pass.

Compliance

Not re-verified in this pass.

Features

  • non-final and final tokens
  • semantic endpoint detection (<end> token, tunable sensitivity and max delay)
  • manual finalize (<fin> token)
  • speaker separation (diarization) included
  • language identification included
  • context / custom vocabulary
  • realtime translation
  • keepalive control message
  • temporary API keys for browsers

Pricing

WhatPriceUnitNotes
Real-time input audio$2.00per 1M audio tokensAbout 30,000 audio tokens per hour of audio.
Real-time input text (context)$4.00per 1M tokens
Real-time output text$4.00per 1M tokensAbout 15,000 output tokens per hour of speech.
Real-time effective rate~$0.12per hourVendor calculator; diarization, language ID and formatting included. Async is ~$0.10/hr.
How the per-minute estimate was worked out

~$0.12/hr vendor estimate = $0.002/min; higher with long context prompts or translation output.

Free tier: Not stated on the pricing page.

Source: soniox.com

Setup

  1. Create an API key in the Soniox Console.
  2. Connect to wss://stt-rt.soniox.com/transcribe-websocket with Authorization: Bearer <key> (or a temporary key for browsers).
  3. Send a JSON start message with model, audio_format, sample_rate, num_channels and options.
  4. Stream binary audio; rebuild text from tokens (append is_final tokens, redraw non-final ones).
  5. Send an empty TEXT frame to finish; wait for {"finished": true}.

Endpoint

wss://stt-rt.soniox.com/transcribe-websocket

Authentication

Authorization: Bearer <API_KEY> header (api_key in the start message is deprecated); temporary API keys for client-side use.

Quick start python

import asyncio, json, os, websockets  # pip install websockets>=14

URL = "wss://stt-rt.soniox.com/transcribe-websocket"
CFG = {"model": "stt-rt-v5", "audio_format": "pcm_s16le", "sample_rate": 16000,
       "num_channels": 1, "enable_endpoint_detection": True}

async def main():
    hdr = {"Authorization": f"Bearer {os.environ['SONIOX_API_KEY']}"}
    async with websockets.connect(URL, additional_headers=hdr) as ws:
        await ws.send(json.dumps(CFG))
        async def send():
            with open("audio_16k_mono.raw", "rb") as f:
                while chunk := f.read(3200):  # 100 ms
                    await ws.send(chunk)
                    await asyncio.sleep(0.1)
            await ws.send("")  # empty TEXT frame ends the stream
        asyncio.create_task(send())
        final = ""
        async for msg in ws:
            d = json.loads(msg)
            if d.get("error_code"):
                print(d); break
            toks = d.get("tokens", [])
            final += "".join(t["text"] for t in toks if t["is_final"])
            partial = "".join(t["text"] for t in toks if not t["is_final"])
            print(final + " | " + partial)
            if d.get("finished"):
                break

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Default concurrency is only 10

New accounts get 10 concurrent streams and 100 requests per minute. Raise it in the console well before a launch.

Hard 300-minute stream cap

Each session ends at 300 minutes and Soniox says this cannot be increased; long broadcasts must roll over to a new session.

Empty binary frame does not end the stream

Only an empty TEXT frame finishes the session; an empty binary frame is treated as an empty audio chunk and the session stays open.

Token-based bill is an estimate

Pricing is per audio, context and output token; the ~$0.12/hr figure assumes typical speech density. Large context prompts and translation output add text tokens.

Model auto-upgrades via aliases

stt-rt-v4 was silently routed to stt-rt-v5 after 2026-06-30. If you need stable behaviour for evaluation, pin the exact model and watch the changelog.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • 100 requests per minute, 10 concurrent streams by default (raisable in console).
  • Each stream capped at 300 minutes; this cap cannot be raised.
  • Open-connection limit is a multiple of the concurrent-request limit, including idle sockets.

Models and products

NameStatusNotes
stt-rt-v5Active (released 2026-06-16)Current realtime model. Alias stt-rt-v4 now points here; v4 removed 2026-06-30 and auto-routed to v5.
stt-async-v5ActiveFile/async counterpart, same accuracy and features per pricing page.

Docs and sources

Docs

Sources used

Not fully verified

Free credit, regions and compliance.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?