GA Kyutai

Unmute

An open, low-latency cascaded voice stack: Kyutai streaming STT plus any OpenAI-compatible text LLM plus Kyutai streaming TTS. The practical way to give a smart self-hosted LLM a natural real-time voice.

Est. per minuten/a

Overview

Best for: Self-hosting a voice front-end for your own (possibly private) LLM with good latency.

At a glance

TypeVoice-to-voice
VRAM GB16
CPU okNo
LicenceCC-BY-4.0
CommercialYes
Full duplexNo
StreamingYes
Latency ms750
Languages2

Cascaded stack. 750 ms is TTS latency on one L40S, about 450 ms with services on separate GPUs. Code MIT, STT/TTS weights CC-BY-4.0; LLM licence depends on your choice. English and French speech.

Audio in

Browser mic via the bundled frontend

Audio out

Streamed TTS audio

Languages

English and French for STT/TTS (per model names); LLM language depends on your model.

Voices

Kyutai TTS voice set (kyutai/tts-voices).

Latency

Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs.

Regions

Self-hosted.

Compliance

Self-hosted; your responsibility.

Hardware

CUDA GPU with at least 16 GB VRAM for the default setup; three or more GPUs recommended to split STT, TTS and LLM for lower latency.

Licence

Unmute code MIT; Kyutai STT/TTS weights CC-BY-4.0; LLM licence depends on your choice.

Features

  • plug in any OpenAI-compatible LLM
  • semantic voice activity / turn-taking
  • voice selection
  • Docker Compose and Swarm deployment

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: github.com

Setup

  1. Linux x86_64 with a CUDA GPU (16 GB is enough for the default Gemma 3 1B setup).
  2. Set your Hugging Face token.
  3. Run docker compose; optionally point KYUTAI_LLM_URL at your own OpenAI-compatible LLM.

Endpoint

Local web app served by the compose stack

Authentication

None by default.

Quick start python

git clone https://github.com/kyutai-labs/unmute && cd unmute
export HUGGING_FACE_HUB_TOKEN=hf_...
# optional: use your own LLM (Ollama, vLLM, OpenRouter...)
# export KYUTAI_LLM_URL=http://host.docker.internal:11434
# export KYUTAI_LLM_MODEL=llama3.1
# export KYUTAI_LLM_API_KEY=...
docker compose up --build

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Cascaded, not native

Tone and emotion in the user's voice are not passed to the LLM; it trades that for a much smarter brain than Moshi.

No tool calling out of the box

Kyutai suggests wrapping vLLM in your own server to return tool-call results; Unmute itself is tool-unaware.

Platform restrictions

x86_64 Linux only (or WSL); no Mac or ARM servers.

English/French speech only

STT/TTS models cover English and French.

GPU count drives latency

Single-GPU deployments roughly double TTS latency versus split GPUs.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • No built-in tool calling (wrap your LLM server to add it)
  • Linux x86_64 only (Windows via WSL); no aarch64 or native Mac

Models and products

NameStatusNotes
kyutai/stt-1b-en_fr, kyutai/stt-2.6b-enGAStreaming STT.
kyutai/tts-1.6b-en_frGAStreaming TTS.
Any OpenAI-compatible LLM (default Gemma 3 1B; vLLM, Ollama, OpenRouter)GA

Docs and sources

Docs

Sources used

Not fully verified

Exact environment variable values for Ollama in the snippet; full language list.

Similar open models APIs

Spotted a wrong price or a dead link?