GA Kyutai

Kyutai TTS (Delayed Streams Modeling)

Streaming-native open TTS that starts speaking before the full text is known, served by a production Rust WebSocket server (the engine behind Kyutai's Unmute). CC-BY-4.0 weights.

Est. per minuten/a

Overview

Best for: Self-hosted agents that need genuine text-in streaming without sentence buffering.

At a glance

TypeText-to-speech
Params B1.6
CPU okNo
LicenceCC-BY-4.0
CommercialYes
StreamingYes
Languages2

True text-input streaming. Values for tts-1.6b-en_fr (also 0.75B English). Each voice has its own licence. Rust server needs an NVIDIA GPU.

Audio in

Streamed text

Audio out

Streamed audio (Mimi codec)

Languages

English, French

Voices

Voice embeddings from the kyutai/tts-voices repository (each voice has its own licence)

Latency

No exact TTS TTFB on the README; the same Rust server is used for Unmute.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

NVIDIA GPU for the Rust server (L40S class used for Unmute)

Licence

Weights CC-BY-4.0 (attribution required); voices have per-voice licences

Features

  • true text-input streaming
  • Rust WebSocket server (moshi-server)
  • batching

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. Clone kyutai-labs/delayed-streams-modeling.
  2. For production install the Rust server (moshi-server) and run: moshi-server worker --config configs/config-tts.toml.
  3. Stream text with scripts/tts_rust_server.py or your own WebSocket client.

Endpoint

Your moshi-server WebSocket

Authentication

Your own

Quick start python

# Quick local test with the PyTorch streaming script from the repo:
#   echo "Hey, how are you?" | python scripts/tts_pytorch_streaming.py audio_output.wav
# Production: Rust server over WebSocket
#   moshi-server worker --config configs/config-tts.toml
#   echo "Hey, how are you?" | python scripts/tts_rust_server.py - -
import subprocess
subprocess.run('echo "Hello from Kyutai." | python scripts/tts_pytorch_streaming.py out.wav',
               shell=True, check=True)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Attribution required

CC-BY-4.0 allows commercial use but requires attribution; check each voice's separate licence in kyutai/tts-voices.

Ops complexity

Production use means running the Rust moshi-server with CUDA; not a pip-install-and-go experience.

Two languages only

English and French.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • English/French only

Models and products

NameStatusNotes
kyutai/tts-1.6b-en_frReleased 2025-06English and French.
kyutai/tts-0.75b-en-publicReleased 2025-07English.

Docs and sources

Docs

Sources used

Not fully verified

Latency and per-GPU concurrency for TTS specifically.

Similar open models APIs

Spotted a wrong price or a dead link?