Kyutai STT (Delayed Streams Modeling)
Open streaming STT models built for realtime: a 1B English/French model with 0.5 s delay and semantic VAD, and a 2.6B English model with 2.5 s delay. Served by a Rust server, PyTorch or MLX.
Overview
Best for: Self-hosted English/French realtime captions or voice agents that want built-in semantic VAD.
At a glance
Values for stt-1b-en_fr (0.5 s delay); stt-2.6b-en is English only with 2.5 s delay. Code Apache-2.0, weights CC-BY-4.0. Apple Silicon via MLX.
Handled by the provided scripts (mic or file); server streams PCM over WebSocket.
n/a
English and French (1B) or English (2.6B).
Model-defined delay: 0.5 s (1B) or 2.5 s (2.6B) per README.
Self-hosted.
Self-hosted.
NVIDIA GPU for moshi-server (README cites H100 and L40S batch sizes); Apple Silicon via MLX for local use.
Code Apache-2.0; model weights CC-BY-4.0.
Features
- word-level timestamps
- semantic VAD (1B)
- batched serving (README: an H100 handles hundreds of streams)
- Apple Silicon via MLX
Pricing
| What | Price | Unit |
|---|---|---|
| Open weights | $0 |
Self-host compute only.
Free tier: Open weights.
Source: github.com
Setup
- Quick test on a file: uvx --with moshi python -m moshi.run_inference --hf-repo kyutai/stt-2.6b-en audio.mp3
- Production: cargo install --features cuda moshi-server, then moshi-server worker --config configs/config-stt-en_fr-hf.toml
- Stream the mic to the server with scripts/stt_from_mic_rust_server.py (or stt_from_mic_mlx.py on a Mac).
Endpoint
ws://localhost:8080 (moshi-server default; check config)
Authentication
Configurable API key in moshi-server config (check repo).
Quick start bash
git clone https://github.com/kyutai-labs/delayed-streams-modeling && cd delayed-streams-modeling
cargo install --features cuda moshi-server
moshi-server worker --config configs/config-stt-en_fr-hf.toml &
uv run scripts/stt_from_mic_rust_server.py # prints words live from your microphone
# Mac without a server: uv run scripts/stt_from_mic_mlx.py
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Model weights are CC-BY-4.0
Code is Apache-2.0 but the model weights on Hugging Face are CC-BY-4.0, which requires attribution in your product.
Only English and French
No other languages are supported by the released STT models.
Repo activity slowed
The GitHub repo's last push was January 2026; expect community support rather than a vendor roadmap.
Rust server build
The production server is a Rust crate built with CUDA features; plan build time and CUDA toolchain on the host.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Fixed algorithmic delay per model; 2.6B is too slow-to-final for snappy voice agents.
Models and products
| Name | Status |
|---|---|
| kyutai/stt-1b-en_fr | Released |
| kyutai/stt-2.6b-en | Released |
Docs and sources
Docs
Sources used
- raw.githubusercontent.com/kyutai-labs/delayed-streams-modeling/main/README.md
- api.github.com/repos/kyutai-labs/delayed-streams-modeling
- huggingface.co/api/models/kyutai/stt-1b-en_fr
- huggingface.co/api/models/kyutai/stt-2.6b-en
Default server port/auth and exact streams-per-GPU figure (README text truncated in fetch).