GA Kyutai

Moshi

The original open full-duplex speech-to-speech model: it listens and talks at the same time with about 200 ms latency. Great for research and natural-sounding chit-chat demos; too small and knowledge-light for most business agents.

Est. per minuten/a
1 high-severity warning

Overview

Best for: Research, latency benchmarks, and natural-feeling chit-chat demos on a single GPU or Mac.

At a glance

TypeVoice-to-voice
Params B7.7
VRAM GB24
CPU okNo
LicenceCC-BY-4.0
CommercialYes
Full duplexYes
StreamingYes
Latency ms200
Languages1

200 ms on an L4 GPU (160 ms theoretical). 24 GB VRAM for PyTorch; MLX quantized builds run on Apple Silicon. Attribution required.

Audio in

24 kHz via the Mimi codec (12.5 Hz frames, 80 ms)

Audio out

24 kHz Mimi-decoded speech

Languages

English (not stated on the repo page; the released models were trained for English).

Voices

Two fixed voices: Moshiko (male) and Moshika (female).

Latency

Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU.

Regions

Self-hosted anywhere.

Compliance

Self-hosted; compliance is your responsibility.

Hardware

PyTorch: NVIDIA GPU with about 24 GB VRAM. MLX q4/q8 tested on MacBook Pro M3. Rust backend needs CUDA or Metal.

Licence

Weights CC-BY-4.0; Python code MIT; Rust backend Apache-2.0.

Features

  • full duplex (overlapping speech, backchannels)
  • barge-in
  • streaming Mimi codec
  • PyTorch, MLX (Apple Silicon) and Rust/Candle backends

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: github.com

Setup

  1. Linux with an NVIDIA GPU (24 GB for PyTorch) or a Mac with Apple Silicon (MLX).
  2. pip install moshi (or moshi_mlx on Mac).
  3. Start the server and open the web UI it serves.

Endpoint

Local: https://localhost:8998 (web UI and WebSocket served by moshi.server)

Authentication

None by default; put it behind your own auth/proxy.

Quick start python

# NVIDIA GPU, PyTorch (about 24 GB VRAM)
pip install -U moshi
python -m moshi.server --hf-repo kyutai/moshika-pytorch-bf16

# Apple Silicon (MLX, 4-bit)
pip install -U moshi_mlx
python -m moshi_mlx.local -q 4 --hf-repo kyutai/moshika-mlx-q4

# Rust/Candle server (CUDA)
# cargo run --features cuda --bin moshi-backend -r -- --config moshi-backend/config.json standalone

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Not an agent brain

The 7B backbone has thin world knowledge and no function calling; it is a conversational demo model, not a support agent.

English only, two voices

No voice selection beyond the two fine-tunes and no multilingual support in the main release.

24 GB VRAM for PyTorch

PyTorch quantization is limited (q8 marked experimental); smaller GPUs need the MLX or Candle builds.

Attribution required

CC-BY-4.0 weights allow commercial use but require attribution to Kyutai.

No auth or multi-tenant serving

The bundled server is single-user oriented; batching and auth are your job.

Research cadence

New checkpoints (RAG, RL 'seamless') appear on Hugging Face without stable product docs.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Small 7B text backbone: limited knowledge and reasoning
  • No built-in tool/function calling
  • Context length not stated on the repo

Models and products

NameStatusNotes
kyutai/moshiko-pytorch-bf16 (male voice), kyutai/moshika-pytorch-bf16 (female voice)GAAbout 7.7B params; also q8 PyTorch (experimental), MLX q4/q8/bf16 and Candle q8/bf16 variants.
kyutai/moshika-rag-pytorch-bf16, kyutai/moshika-rl-seamlessPreviewNewer 2026 research checkpoints on Hugging Face (RAG and RL fine-tunes); details not verified.

Docs and sources

Docs

Sources used

Not fully verified

Exact MLX repo name in the snippet; context limit; details of the 2026 RAG/RL checkpoints.

Similar open models APIs

Spotted a wrong price or a dead link?