Unmute
An open, low-latency cascaded voice stack: Kyutai streaming STT plus any OpenAI-compatible text LLM plus Kyutai streaming TTS. The practical way to give a smart self-hosted LLM a natural real-time voice.
Overview
Best for: Self-hosting a voice front-end for your own (possibly private) LLM with good latency.
At a glance
Cascaded stack. 750 ms is TTS latency on one L40S, about 450 ms with services on separate GPUs. Code MIT, STT/TTS weights CC-BY-4.0; LLM licence depends on your choice. English and French speech.
Browser mic via the bundled frontend
Streamed TTS audio
English and French for STT/TTS (per model names); LLM language depends on your model.
Kyutai TTS voice set (kyutai/tts-voices).
Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs.
Self-hosted.
Self-hosted; your responsibility.
CUDA GPU with at least 16 GB VRAM for the default setup; three or more GPUs recommended to split STT, TTS and LLM for lower latency.
Unmute code MIT; Kyutai STT/TTS weights CC-BY-4.0; LLM licence depends on your choice.
Features
- plug in any OpenAI-compatible LLM
- semantic voice activity / turn-taking
- voice selection
- Docker Compose and Swarm deployment
Pricing
No public price list.
No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint.
n/a (self-hosted)
Free tier: Open weights
Source: github.com
Setup
- Linux x86_64 with a CUDA GPU (16 GB is enough for the default Gemma 3 1B setup).
- Set your Hugging Face token.
- Run docker compose; optionally point KYUTAI_LLM_URL at your own OpenAI-compatible LLM.
Endpoint
Local web app served by the compose stack
Authentication
None by default.
Quick start python
git clone https://github.com/kyutai-labs/unmute && cd unmute
export HUGGING_FACE_HUB_TOKEN=hf_...
# optional: use your own LLM (Ollama, vLLM, OpenRouter...)
# export KYUTAI_LLM_URL=http://host.docker.internal:11434
# export KYUTAI_LLM_MODEL=llama3.1
# export KYUTAI_LLM_API_KEY=...
docker compose up --build
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Cascaded, not native
Tone and emotion in the user's voice are not passed to the LLM; it trades that for a much smarter brain than Moshi.
No tool calling out of the box
Kyutai suggests wrapping vLLM in your own server to return tool-call results; Unmute itself is tool-unaware.
Platform restrictions
x86_64 Linux only (or WSL); no Mac or ARM servers.
English/French speech only
STT/TTS models cover English and French.
GPU count drives latency
Single-GPU deployments roughly double TTS latency versus split GPUs.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- No built-in tool calling (wrap your LLM server to add it)
- Linux x86_64 only (Windows via WSL); no aarch64 or native Mac
Models and products
| Name | Status |
|---|---|
| kyutai/stt-1b-en_fr, kyutai/stt-2.6b-en | GA |
| kyutai/tts-1.6b-en_fr | GA |
| Any OpenAI-compatible LLM (default Gemma 3 1B; vLLM, Ollama, OpenRouter) | GA |
Docs and sources
Docs
Sources used
Exact environment variable values for Ollama in the snippet; full language list.