Speechmatics Realtime and Agent STT
Long-running WebSocket realtime transcription with Enhanced/Standard models, realtime diarization and big custom dictionaries; a separate Agent STT endpoint (Linden 1) gives turn-based, speaker-labelled output for voice agents.
Overview
Best for: Long sessions (up to 48 h), broadcast and meeting captions with realtime diarization and large custom vocabularies; EU-hosted processing.
At a glance
Languages: vendor markets 50+, count not verified. Price is Realtime Standard (default); Enhanced $0.80/hr, Linden 1 agent model $0.30/hr. Concurrency is the Pro plan (Free 2). Session max 48 h. Realtime code-switching only via Melia 1, which is a preview. Latency tunable via max_delay (0.7 s in examples), no vendor claim.
raw pcm_f32le, pcm_s16le or mulaw with explicit sample_rate; or 'file' type for wav, mp3, aac, ogg, mpeg, amr, m4a, mp4, flac.
n/a
Language packs per session (vendor markets 50+ languages; count not re-verified). Enhanced/Standard need a selected language; Melia 1 handles code-switching.
Configurable via max_delay (docs examples use 0.7 s); no other vendor latency figure captured.
EU, US, AUS listed for models; global endpoint routes to nearest region; regional endpoints for residency.
Docs state Realtime SaaS does not store audio, transcripts or configuration. On-prem containers and Kubernetes deployments documented. Certifications not re-verified.
Features
- partials (enable_partials)
- max_delay latency control
- realtime diarization (speaker and channel)
- speaker identification
- custom dictionary up to 20,000 items
- turn detection / force end of utterance
- translation add-on
- Agent STT with turn events and speaker-labelled segments
Pricing
| What | Price | Unit |
|---|---|---|
| Realtime Standard | $0.45 | per hour |
| Realtime Enhanced | $0.80 | per hour |
| Linden 1 (Agent STT) | $0.30 | per hour |
| Melia 1 | $0.40 | per hour |
| Translation add-on | $0.65 | per hour |
Linden 1 with training discount ($0.20/hr) up to Enhanced list ($0.80/hr). Subscriptions/credit packs cut up to 25%/20%.
Free tier: $100 in credits at sign-up, no card (pricing page).
Source: speechmatics.com
Setup
- Create an API key in the Speechmatics portal.
- Connect to wss://eu.rt.speechmatics.com/v2 (or global.rt.speechmatics.com/v2) with Authorization: Bearer <key>, or a short-lived JWT for browsers.
- Send StartRecognition JSON with audio_format and transcription_config, wait for RecognitionStarted.
- Send binary audio (AddAudio), read AddPartialTranscript / AddTranscript, finish with EndOfStream {last_seq_no}.
- For voice agents use wss://global.rt.speechmatics.com/v2/agent (Agent STT).
Endpoint
wss://eu.rt.speechmatics.com/v2 ; wss://global.rt.speechmatics.com/v2 ; Agent STT: wss://global.rt.speechmatics.com/v2/agent
Authentication
Authorization: Bearer <API_KEY>; browser clients use a temporary JWT (createSpeechmaticsJWT with type rt and a ttl), passed as ?jwt=.
Quick start python
import asyncio, json, os, websockets # pip install websockets>=14
URL = "wss://eu.rt.speechmatics.com/v2"
START = {"message": "StartRecognition",
"audio_format": {"type": "raw", "encoding": "pcm_s16le", "sample_rate": 16000},
"transcription_config": {"language": "en", "model": "enhanced",
"enable_partials": True, "max_delay": 0.7}}
async def main():
hdr = {"Authorization": f"Bearer {os.environ['SPEECHMATICS_API_KEY']}"}
async with websockets.connect(URL, additional_headers=hdr) as ws:
await ws.send(json.dumps(START))
async def send():
n = 0
with open("audio_16k_mono.raw", "rb") as f:
while chunk := f.read(3200): # 100 ms
await ws.send(chunk); n += 1
await asyncio.sleep(0.1)
await ws.send(json.dumps({"message": "EndOfStream", "last_seq_no": n}))
async for msg in ws:
d = json.loads(msg)
if d["message"] == "RecognitionStarted":
asyncio.create_task(send())
elif d["message"] in ("AddPartialTranscript", "AddTranscript"):
kind = "partial" if d["message"] == "AddPartialTranscript" else "FINAL"
print(kind, d["metadata"]["transcript"])
elif d["message"] in ("EndOfTranscript", "Error"):
print(d); break
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
3-minute silence kill
A session with no audio and no ping/pong for 3 minutes is ended (and after 1 hour without AddAudio). Keep sending audio or WebSocket pings during holds.
Free tier is 2 concurrent sessions
Free accounts can only run 2 realtime sessions at once; Pro is 50. Load tests on a free key fail early.
Multilingual models are not realtime GA
Melia 1 and Oak 1 are documented as batch; realtime Melia 1 is a preview page. Realtime GA is Enhanced/Standard, which need you to pick a language.
Training discount means data use
The cheaper '-33%' prices come from opting into model training. Check that this is acceptable for your data before picking those rates.
Agent STT is a different API
Linden 1 runs only on the /v2/agent endpoint with its own message schema and is SaaS only; it is not selectable with model= on the normal realtime API.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- Concurrent realtime sessions: Free 2, Pro 50, Enterprise custom.
- Session ends at 48 hours, after 1 hour without AddAudio, or after 3 minutes with no audio or ping/pong.
- Custom dictionary over 20,000 items closes the socket with protocol_error.
Models and products
| Name | Status |
|---|---|
| enhanced | GA |
| standard | GA (default) |
| linden-1 | GA (Agent STT only) |
| melia-1 | Preview for realtime |
| oak-1 | Batch only |
Docs and sources
Docs
Sources used
- speechmatics.com/pricing
- docs.speechmatics.com/speech-to-text/models.md
- docs.speechmatics.com/speech-to-text/realtime/limits.md
- docs.speechmatics.com/speech-to-text/realtime/input.md
- docs.speechmatics.com/speech-to-text/agent-stt.md
Exact language count and whether realtime is billed by session length or audio length were not confirmed.