Moshi
The original open full-duplex speech-to-speech model: it listens and talks at the same time with about 200 ms latency. Great for research and natural-sounding chit-chat demos; too small and knowledge-light for most business agents.
Overview
Best for: Research, latency benchmarks, and natural-feeling chit-chat demos on a single GPU or Mac.
At a glance
200 ms on an L4 GPU (160 ms theoretical). 24 GB VRAM for PyTorch; MLX quantized builds run on Apple Silicon. Attribution required.
24 kHz via the Mimi codec (12.5 Hz frames, 80 ms)
24 kHz Mimi-decoded speech
English (not stated on the repo page; the released models were trained for English).
Two fixed voices: Moshiko (male) and Moshika (female).
Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU.
Self-hosted anywhere.
Self-hosted; compliance is your responsibility.
PyTorch: NVIDIA GPU with about 24 GB VRAM. MLX q4/q8 tested on MacBook Pro M3. Rust backend needs CUDA or Metal.
Weights CC-BY-4.0; Python code MIT; Rust backend Apache-2.0.
Features
- full duplex (overlapping speech, backchannels)
- barge-in
- streaming Mimi codec
- PyTorch, MLX (Apple Silicon) and Rust/Candle backends
Pricing
No public price list.
No licence fee; you pay for GPU time.
n/a (self-hosted)
Free tier: Open weights
Source: github.com
Setup
- Linux with an NVIDIA GPU (24 GB for PyTorch) or a Mac with Apple Silicon (MLX).
- pip install moshi (or moshi_mlx on Mac).
- Start the server and open the web UI it serves.
Endpoint
Local: https://localhost:8998 (web UI and WebSocket served by moshi.server)
Authentication
None by default; put it behind your own auth/proxy.
Quick start python
# NVIDIA GPU, PyTorch (about 24 GB VRAM)
pip install -U moshi
python -m moshi.server --hf-repo kyutai/moshika-pytorch-bf16
# Apple Silicon (MLX, 4-bit)
pip install -U moshi_mlx
python -m moshi_mlx.local -q 4 --hf-repo kyutai/moshika-mlx-q4
# Rust/Candle server (CUDA)
# cargo run --features cuda --bin moshi-backend -r -- --config moshi-backend/config.json standalone
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Not an agent brain
The 7B backbone has thin world knowledge and no function calling; it is a conversational demo model, not a support agent.
English only, two voices
No voice selection beyond the two fine-tunes and no multilingual support in the main release.
24 GB VRAM for PyTorch
PyTorch quantization is limited (q8 marked experimental); smaller GPUs need the MLX or Candle builds.
Attribution required
CC-BY-4.0 weights allow commercial use but require attribution to Kyutai.
No auth or multi-tenant serving
The bundled server is single-user oriented; batching and auth are your job.
Research cadence
New checkpoints (RAG, RL 'seamless') appear on Hugging Face without stable product docs.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Small 7B text backbone: limited knowledge and reasoning
- No built-in tool/function calling
- Context length not stated on the repo
Models and products
| Name | Status |
|---|---|
| kyutai/moshiko-pytorch-bf16 (male voice), kyutai/moshika-pytorch-bf16 (female voice) | GA |
| kyutai/moshika-rag-pytorch-bf16, kyutai/moshika-rl-seamless | Preview |
Docs and sources
Docs
Sources used
- github.com/kyutai-labs/moshi
- huggingface.co/api/models/kyutai/moshiko-pytorch-bf16
- huggingface.co/api/models?author=kyutai
Exact MLX repo name in the snippet; context limit; details of the 2026 RAG/RL checkpoints.