GA Mistral AI

Mistral Voxtral Mini Transcribe Realtime

Hosted realtime transcription over WebSocket at $0.006/min, and the same 4B model released as Apache-2.0 open weights for self-hosting.

Est. per minute$0.006

Overview

Best for: EU-vendor realtime STT with an escape hatch to run the identical model on your own GPUs.

At a glance

$/hour$0.36
Live speakersNo
PartialsYes
Latency ms200
Languages13
WebRTCNo
WebSocketYes
gRPCNo
Self-hostYes
Open weightsYes

Latency configurable down to sub-200 ms (vendor). $0.006/min from the launch announcement, not a live price table. Open weights Apache 2.0 (Voxtral-Mini-4B-Realtime). Realtime cannot be combined with diarization.

Audio in

AudioFormat(encoding='pcm_s16le', sample_rate=16000) in docs examples; mic example sends 480 ms chunks.

Audio out

n/a

Languages

13 languages (vendor).

Latency

Vendor claim: configurable latency down to sub-200 ms; press reports ~1-2% error rate at 480 ms delay (secondary source).

Regions

Mistral API (EU company); self-host anywhere.

Compliance

Not re-verified. Self-hosting keeps audio fully in your infrastructure.

Hardware

4B-parameter model; GPU recommended for realtime serving. Mistral says the footprint can run on edge devices (not benchmarked here).

Licence

Apache 2.0 (open weights).

Features

  • streaming text deltas
  • configurable latency
  • short-lived rt_ tokens for browsers
  • open weights for on-prem

Pricing

WhatPriceUnitNotes
Voxtral Realtime API$0.006per minuteFrom Mistral's Voxtral Transcribe 2 announcement; batch Voxtral is $0.003/min.
Self-hosted open weightsfreeApache 2.0; you pay for GPUs.
How the per-minute estimate was worked out

Flat announced rate; not confirmed on a live pricing table in this pass.

Free tier: Not stated.

Source: mistral.ai

Setup

  1. pip install "mistralai[realtime]>=2" (V2 SDK).
  2. Stream PCM to client.audio.realtime.transcribe_stream with model voxtral-mini-transcribe-realtime-2602.
  3. For browsers: server calls POST https://api.mistral.ai/v1/client/sessions {purpose:'realtime', model} to mint an rt_ token; browser opens the WebSocket with subprotocols ['realtime', token].
  4. Self-host: download the HF weights and serve with an engine that supports the realtime model (check model card).

Endpoint

wss://api.mistral.ai/v1/audio/transcriptions/realtime?model=voxtral-mini-transcribe-realtime-2602

Authentication

Authorization: Bearer <key> server side; rt_ short-lived token via Sec-WebSocket-Protocol in browsers.

Quick start python

import asyncio, os  # pip install "mistralai[realtime]>=2"
from mistralai.client import Mistral
from mistralai.client.models import (AudioFormat, TranscriptionStreamDone,
                                     TranscriptionStreamTextDelta)

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
fmt = AudioFormat(encoding="pcm_s16le", sample_rate=16000)

async def audio_chunks():
    with open("audio_16k_mono.raw", "rb") as f:
        while chunk := f.read(15360):  # 480 ms
            yield chunk
            await asyncio.sleep(0.48)

async def main():
    async for ev in client.audio.realtime.transcribe_stream(
        audio_stream=audio_chunks(),
        model="voxtral-mini-transcribe-realtime-2602",
        audio_format=fmt,
    ):
        if isinstance(ev, TranscriptionStreamTextDelta):
            print(ev.text, end="", flush=True)
        elif isinstance(ev, TranscriptionStreamDone):
            print("\n[done]")

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No diarization in realtime

Mistral's docs state realtime is not compatible with the diarize parameter; use one or the other.

Only 13 languages

Strong performance is claimed for 13 languages, far fewer than Soniox, Gladia or Google. Check your language first.

SDK v1 cannot do realtime

Realtime transcription is only in the V2 Python SDK (mistralai>=2); v1 code examples will not work.

Short token lifetime

Browser rt_ tokens last about 900 s; mint right before the user starts talking, not at page load.

Price from announcement

The $0.006/min rate comes from the launch post; confirm on the console pricing table before committing.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • rt_ client tokens expire after about 900 seconds; mint them just before connecting.
  • Python realtime needs SDK v2 (mistralai>=2, extra [realtime]); not available in v1 SDK.

Models and products

NameStatusNotes
voxtral-mini-transcribe-realtime-2602GA (API)Only realtime-capable model; not compatible with the diarize parameter.
mistralai/Voxtral-Mini-4B-Realtime-2602Open weightsApache 2.0 on Hugging Face; 4B parameters.

Docs and sources

Docs

Sources used

Not fully verified

Price on a live pricing table, language list, concurrency limits, and the recommended self-host serving stack.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?