Kyutai TTS (Delayed Streams Modeling)
Streaming-native open TTS that starts speaking before the full text is known, served by a production Rust WebSocket server (the engine behind Kyutai's Unmute). CC-BY-4.0 weights.
Overview
Best for: Self-hosted agents that need genuine text-in streaming without sentence buffering.
At a glance
True text-input streaming. Values for tts-1.6b-en_fr (also 0.75B English). Each voice has its own licence. Rust server needs an NVIDIA GPU.
Streamed text
Streamed audio (Mimi codec)
English, French
Voice embeddings from the kyutai/tts-voices repository (each voice has its own licence)
No exact TTS TTFB on the README; the same Rust server is used for Unmute.
Wherever you deploy it
Your own deployment; no vendor data processing
NVIDIA GPU for the Rust server (L40S class used for Unmute)
Weights CC-BY-4.0 (attribution required); voices have per-voice licences
Features
- true text-input streaming
- Rust WebSocket server (moshi-server)
- batching
Pricing
| What | Price | Unit |
|---|---|---|
| Weights | $0 |
Self-hosted: cost is your GPU/CPU time, not per character
Free tier: Open weights
Source: huggingface.co
Setup
- Clone kyutai-labs/delayed-streams-modeling.
- For production install the Rust server (moshi-server) and run: moshi-server worker --config configs/config-tts.toml.
- Stream text with scripts/tts_rust_server.py or your own WebSocket client.
Endpoint
Your moshi-server WebSocket
Authentication
Your own
Quick start python
# Quick local test with the PyTorch streaming script from the repo:
# echo "Hey, how are you?" | python scripts/tts_pytorch_streaming.py audio_output.wav
# Production: Rust server over WebSocket
# moshi-server worker --config configs/config-tts.toml
# echo "Hey, how are you?" | python scripts/tts_rust_server.py - -
import subprocess
subprocess.run('echo "Hello from Kyutai." | python scripts/tts_pytorch_streaming.py out.wav',
shell=True, check=True)
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Attribution required
CC-BY-4.0 allows commercial use but requires attribution; check each voice's separate licence in kyutai/tts-voices.
Ops complexity
Production use means running the Rust moshi-server with CUDA; not a pip-install-and-go experience.
Two languages only
English and French.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- English/French only
Models and products
| Name | Status |
|---|---|
| kyutai/tts-1.6b-en_fr | Released 2025-06 |
| kyutai/tts-0.75b-en-public | Released 2025-07 |
Docs and sources
Docs
Sources used
Latency and per-GPU concurrency for TTS specifically.