GA Kyutai

Kyutai STT (Delayed Streams Modeling)

Open streaming STT models built for realtime: a 1B English/French model with 0.5 s delay and semantic VAD, and a 2.6B English model with 2.5 s delay. Served by a Rust server, PyTorch or MLX.

Est. per minute$0

Overview

Best for: Self-hosted English/French realtime captions or voice agents that want built-in semantic VAD.

At a glance

TypeSpeech-to-text
Params B1
LicenceCC-BY-4.0
CommercialYes
StreamingYes
Latency ms500
Languages2

Values for stt-1b-en_fr (0.5 s delay); stt-2.6b-en is English only with 2.5 s delay. Code Apache-2.0, weights CC-BY-4.0. Apple Silicon via MLX.

Audio in

Handled by the provided scripts (mic or file); server streams PCM over WebSocket.

Audio out

n/a

Languages

English and French (1B) or English (2.6B).

Latency

Model-defined delay: 0.5 s (1B) or 2.5 s (2.6B) per README.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

NVIDIA GPU for moshi-server (README cites H100 and L40S batch sizes); Apple Silicon via MLX for local use.

Licence

Code Apache-2.0; model weights CC-BY-4.0.

Features

  • word-level timestamps
  • semantic VAD (1B)
  • batched serving (README: an H100 handles hundreds of streams)
  • Apple Silicon via MLX

Pricing

WhatPriceUnitNotes
Open weights$0Compute only.
How the per-minute estimate was worked out

Self-host compute only.

Free tier: Open weights.

Source: github.com

Setup

  1. Quick test on a file: uvx --with moshi python -m moshi.run_inference --hf-repo kyutai/stt-2.6b-en audio.mp3
  2. Production: cargo install --features cuda moshi-server, then moshi-server worker --config configs/config-stt-en_fr-hf.toml
  3. Stream the mic to the server with scripts/stt_from_mic_rust_server.py (or stt_from_mic_mlx.py on a Mac).

Endpoint

ws://localhost:8080 (moshi-server default; check config)

Authentication

Configurable API key in moshi-server config (check repo).

Quick start bash

git clone https://github.com/kyutai-labs/delayed-streams-modeling && cd delayed-streams-modeling
cargo install --features cuda moshi-server
moshi-server worker --config configs/config-stt-en_fr-hf.toml &
uv run scripts/stt_from_mic_rust_server.py   # prints words live from your microphone
# Mac without a server: uv run scripts/stt_from_mic_mlx.py

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Model weights are CC-BY-4.0

Code is Apache-2.0 but the model weights on Hugging Face are CC-BY-4.0, which requires attribution in your product.

Only English and French

No other languages are supported by the released STT models.

Repo activity slowed

The GitHub repo's last push was January 2026; expect community support rather than a vendor roadmap.

Rust server build

The production server is a Rust crate built with CUDA features; plan build time and CUDA toolchain on the host.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Fixed algorithmic delay per model; 2.6B is too slow-to-final for snappy voice agents.

Models and products

NameStatusNotes
kyutai/stt-1b-en_frReleasedEnglish + French, ~1B params, 0.5 s delay, semantic VAD.
kyutai/stt-2.6b-enReleasedEnglish only, ~2.6B params, 2.5 s delay, higher accuracy.

Docs and sources

Docs

Sources used

Not fully verified

Default server port/auth and exact streams-per-GPU figure (README text truncated in fetch).

Similar open models APIs

Spotted a wrong price or a dead link?