Mistral Voxtral Mini Transcribe Realtime
Hosted realtime transcription over WebSocket at $0.006/min, and the same 4B model released as Apache-2.0 open weights for self-hosting.
Overview
Best for: EU-vendor realtime STT with an escape hatch to run the identical model on your own GPUs.
At a glance
Latency configurable down to sub-200 ms (vendor). $0.006/min from the launch announcement, not a live price table. Open weights Apache 2.0 (Voxtral-Mini-4B-Realtime). Realtime cannot be combined with diarization.
AudioFormat(encoding='pcm_s16le', sample_rate=16000) in docs examples; mic example sends 480 ms chunks.
n/a
13 languages (vendor).
Vendor claim: configurable latency down to sub-200 ms; press reports ~1-2% error rate at 480 ms delay (secondary source).
Mistral API (EU company); self-host anywhere.
Not re-verified. Self-hosting keeps audio fully in your infrastructure.
4B-parameter model; GPU recommended for realtime serving. Mistral says the footprint can run on edge devices (not benchmarked here).
Apache 2.0 (open weights).
Features
- streaming text deltas
- configurable latency
- short-lived rt_ tokens for browsers
- open weights for on-prem
Pricing
| What | Price | Unit |
|---|---|---|
| Voxtral Realtime API | $0.006 | per minute |
| Self-hosted open weights | free |
Flat announced rate; not confirmed on a live pricing table in this pass.
Free tier: Not stated.
Source: mistral.ai
Setup
- pip install "mistralai[realtime]>=2" (V2 SDK).
- Stream PCM to client.audio.realtime.transcribe_stream with model voxtral-mini-transcribe-realtime-2602.
- For browsers: server calls POST https://api.mistral.ai/v1/client/sessions {purpose:'realtime', model} to mint an rt_ token; browser opens the WebSocket with subprotocols ['realtime', token].
- Self-host: download the HF weights and serve with an engine that supports the realtime model (check model card).
Endpoint
wss://api.mistral.ai/v1/audio/transcriptions/realtime?model=voxtral-mini-transcribe-realtime-2602
Authentication
Authorization: Bearer <key> server side; rt_ short-lived token via Sec-WebSocket-Protocol in browsers.
Quick start python
import asyncio, os # pip install "mistralai[realtime]>=2"
from mistralai.client import Mistral
from mistralai.client.models import (AudioFormat, TranscriptionStreamDone,
TranscriptionStreamTextDelta)
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
fmt = AudioFormat(encoding="pcm_s16le", sample_rate=16000)
async def audio_chunks():
with open("audio_16k_mono.raw", "rb") as f:
while chunk := f.read(15360): # 480 ms
yield chunk
await asyncio.sleep(0.48)
async def main():
async for ev in client.audio.realtime.transcribe_stream(
audio_stream=audio_chunks(),
model="voxtral-mini-transcribe-realtime-2602",
audio_format=fmt,
):
if isinstance(ev, TranscriptionStreamTextDelta):
print(ev.text, end="", flush=True)
elif isinstance(ev, TranscriptionStreamDone):
print("\n[done]")
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
No diarization in realtime
Mistral's docs state realtime is not compatible with the diarize parameter; use one or the other.
Only 13 languages
Strong performance is claimed for 13 languages, far fewer than Soniox, Gladia or Google. Check your language first.
SDK v1 cannot do realtime
Realtime transcription is only in the V2 Python SDK (mistralai>=2); v1 code examples will not work.
Short token lifetime
Browser rt_ tokens last about 900 s; mint right before the user starts talking, not at page load.
Price from announcement
The $0.006/min rate comes from the launch post; confirm on the console pricing table before committing.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- rt_ client tokens expire after about 900 seconds; mint them just before connecting.
- Python realtime needs SDK v2 (mistralai>=2, extra [realtime]); not available in v1 SDK.
Models and products
| Name | Status |
|---|---|
| voxtral-mini-transcribe-realtime-2602 | GA (API) |
| mistralai/Voxtral-Mini-4B-Realtime-2602 | Open weights |
Docs and sources
Docs
Sources used
- docs.mistral.ai/studio/audio/speech_to_text/realtime_transcription.md
- docs.mistral.ai/studio/audio/speech_to_text/realtime_transcription/client_auth....
- docs.mistral.ai/capabilities/audio_transcription
- mistral.ai/news/voxtral-transcribe-2
- the-decoder.com/voxtral-transcribe-2-offers-speech-recognition-at-0-003-per-min...
- huggingface.co/api/models/mistralai/Voxtral-Mini-4B-Realtime-2602
Price on a live pricing table, language list, concurrency limits, and the recommended self-host serving stack.