GA Speechmatics

Speechmatics Realtime and Agent STT

Long-running WebSocket realtime transcription with Enhanced/Standard models, realtime diarization and big custom dictionaries; a separate Agent STT endpoint (Linden 1) gives turn-based, speaker-labelled output for voice agents.

Est. per minute$0.0033 - 0.013

Overview

Best for: Long sessions (up to 48 h), broadcast and meeting captions with realtime diarization and large custom vocabularies; EU-hosted processing.

At a glance

$/hour$0.45
Free tierYes
Free credit $$100
Live speakersYes
Turn detectYes
KeytermsYes
PartialsYes
Mixed langsYes
8 kHz phoneYes
Languages50
Max session min2,880
Concurrency50
WebRTCNo
WebSocketYes
gRPCNo
EU dataYes
Self-hostYes
Open weightsNo

Languages: vendor markets 50+, count not verified. Price is Realtime Standard (default); Enhanced $0.80/hr, Linden 1 agent model $0.30/hr. Concurrency is the Pro plan (Free 2). Session max 48 h. Realtime code-switching only via Melia 1, which is a preview. Latency tunable via max_delay (0.7 s in examples), no vendor claim.

Audio in

raw pcm_f32le, pcm_s16le or mulaw with explicit sample_rate; or 'file' type for wav, mp3, aac, ogg, mpeg, amr, m4a, mp4, flac.

Audio out

n/a

Languages

Language packs per session (vendor markets 50+ languages; count not re-verified). Enhanced/Standard need a selected language; Melia 1 handles code-switching.

Latency

Configurable via max_delay (docs examples use 0.7 s); no other vendor latency figure captured.

Regions

EU, US, AUS listed for models; global endpoint routes to nearest region; regional endpoints for residency.

Compliance

Docs state Realtime SaaS does not store audio, transcripts or configuration. On-prem containers and Kubernetes deployments documented. Certifications not re-verified.

Features

  • partials (enable_partials)
  • max_delay latency control
  • realtime diarization (speaker and channel)
  • speaker identification
  • custom dictionary up to 20,000 items
  • turn detection / force end of utterance
  • translation add-on
  • Agent STT with turn events and speaker-labelled segments

Pricing

WhatPriceUnitNotes
Realtime Standard$0.45per hour$0.30/hr with the model-training discount (-33%).
Realtime Enhanced$0.80per hour$0.54/hr with model-training discount.
Linden 1 (Agent STT)$0.30per hour$0.20/hr with model-training discount.
Melia 1$0.40per hourListed price; realtime is preview.
Translation add-on$0.65per hourSupported for realtime.
How the per-minute estimate was worked out

Linden 1 with training discount ($0.20/hr) up to Enhanced list ($0.80/hr). Subscriptions/credit packs cut up to 25%/20%.

Free tier: $100 in credits at sign-up, no card (pricing page).

Source: speechmatics.com

Setup

  1. Create an API key in the Speechmatics portal.
  2. Connect to wss://eu.rt.speechmatics.com/v2 (or global.rt.speechmatics.com/v2) with Authorization: Bearer <key>, or a short-lived JWT for browsers.
  3. Send StartRecognition JSON with audio_format and transcription_config, wait for RecognitionStarted.
  4. Send binary audio (AddAudio), read AddPartialTranscript / AddTranscript, finish with EndOfStream {last_seq_no}.
  5. For voice agents use wss://global.rt.speechmatics.com/v2/agent (Agent STT).

Endpoint

wss://eu.rt.speechmatics.com/v2 ; wss://global.rt.speechmatics.com/v2 ; Agent STT: wss://global.rt.speechmatics.com/v2/agent

Authentication

Authorization: Bearer <API_KEY>; browser clients use a temporary JWT (createSpeechmaticsJWT with type rt and a ttl), passed as ?jwt=.

Quick start python

import asyncio, json, os, websockets  # pip install websockets>=14

URL = "wss://eu.rt.speechmatics.com/v2"
START = {"message": "StartRecognition",
         "audio_format": {"type": "raw", "encoding": "pcm_s16le", "sample_rate": 16000},
         "transcription_config": {"language": "en", "model": "enhanced",
                                  "enable_partials": True, "max_delay": 0.7}}

async def main():
    hdr = {"Authorization": f"Bearer {os.environ['SPEECHMATICS_API_KEY']}"}
    async with websockets.connect(URL, additional_headers=hdr) as ws:
        await ws.send(json.dumps(START))
        async def send():
            n = 0
            with open("audio_16k_mono.raw", "rb") as f:
                while chunk := f.read(3200):  # 100 ms
                    await ws.send(chunk); n += 1
                    await asyncio.sleep(0.1)
            await ws.send(json.dumps({"message": "EndOfStream", "last_seq_no": n}))
        async for msg in ws:
            d = json.loads(msg)
            if d["message"] == "RecognitionStarted":
                asyncio.create_task(send())
            elif d["message"] in ("AddPartialTranscript", "AddTranscript"):
                kind = "partial" if d["message"] == "AddPartialTranscript" else "FINAL"
                print(kind, d["metadata"]["transcript"])
            elif d["message"] in ("EndOfTranscript", "Error"):
                print(d); break

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

3-minute silence kill

A session with no audio and no ping/pong for 3 minutes is ended (and after 1 hour without AddAudio). Keep sending audio or WebSocket pings during holds.

Free tier is 2 concurrent sessions

Free accounts can only run 2 realtime sessions at once; Pro is 50. Load tests on a free key fail early.

Multilingual models are not realtime GA

Melia 1 and Oak 1 are documented as batch; realtime Melia 1 is a preview page. Realtime GA is Enhanced/Standard, which need you to pick a language.

Training discount means data use

The cheaper '-33%' prices come from opting into model training. Check that this is acceptable for your data before picking those rates.

Agent STT is a different API

Linden 1 runs only on the /v2/agent endpoint with its own message schema and is SaaS only; it is not selectable with model= on the normal realtime API.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Concurrent realtime sessions: Free 2, Pro 50, Enterprise custom.
  • Session ends at 48 hours, after 1 hour without AddAudio, or after 3 minutes with no audio or ping/pong.
  • Custom dictionary over 20,000 items closes the socket with protocol_error.

Models and products

NameStatusNotes
enhancedGAHighest accuracy, single-language, medical variant available. Realtime and batch.
standardGA (default)Faster/cheaper; used if model is not set.
linden-1GA (Agent STT only)Voice-agent model on wss://.../v2/agent with turn detection, segmentation and diarization.
melia-1Preview for realtimeMultilingual code-switching (56 languages in realtime preview page); batch is the documented mode.
oak-1Batch onlyMultilingual healthcare model.

Docs and sources

Docs

Sources used

Not fully verified

Exact language count and whether realtime is billed by session length or audio length were not confirmed.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?