Mistral Voxtral TTS
4B-parameter TTS released March 2026 with 20 preset voices in 9 languages and few-second voice adaptation. Hosted on Mistral's /v1/audio/speech; open weights are non-commercial (CC BY-NC 4.0).
Overview
Best for: European-language agents on Mistral; research self-hosting.
At a glance
70 ms is the self-hosted model card at concurrency 1 (552 ms at 32). $16/1M from a search snippet of the pricing page. Open weights are CC BY-NC 4.0 (non-commercial). Cloning means voice adaptation.
Text
Not verified
9: en, fr, es, pt, it, nl, de, ar, hi
20 presets; adapts to new voices from ~3 s reference (secondary source)
Model card (self-hosted, 1 GPU): 70 ms at concurrency 1, 331 ms at 16, 552 ms at 32.
Not verified
Not verified.
Features
- streaming and batch
- voice adaptation
- open weights (non-commercial)
Pricing
| What | Price | Unit |
|---|---|---|
| Voxtral TTS | $0.016 | per 1K characters |
900 chars/min x $16/1M
Free tier: Not verified
Source: mistral.ai
Setup
- Create a Mistral API key.
- POST to https://api.mistral.ai/v1/audio/speech with model voxtral-mini-tts-2603 (check API reference for voice parameter).
- Or self-host the weights with vLLM Omni (non-commercial only).
Endpoint
https://api.mistral.ai/v1/audio/speech
Authentication
Authorization: Bearer <MISTRAL_API_KEY>
Warnings
Weights are non-commercial
Voxtral-4B-TTS-2603 and its reference voices are CC BY-NC 4.0; commercial products must use the paid API or a separate licence.
Latency rises fast with concurrency
Mistral's own table shows 70 ms at 1 stream but 552 ms at 32 streams per GPU.
Price not directly confirmed
The $0.016/1K figure came from a search snippet of Mistral's pricing page; confirm in the console.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Self-host: single GPU with >= 16 GB memory (BF16)
Models and products
| Name | Status |
|---|---|
| voxtral-mini-tts-2603 (API) / mistralai/Voxtral-4B-TTS-2603 (weights) | GA |
Docs and sources
Docs
Sources used
API price (snippet only), output formats, input streaming support, voice parameter name.