GA Kyutai

Kyutai Pocket TTS

Small CPU-only TTS with voice cloning: ~200 ms to first audio chunk and ~6x real time on a MacBook Air M4 using 2 cores. Six languages.

Est. per minuten/a

Overview

Best for: On-device or cheap CPU servers needing cloning in European languages.

At a glance

TypeText-to-speech
CPU okYes
LicenceCC-BY-4.0
CommercialYes
StreamingYes
Latency ms200
Languages6

200 ms to first chunk on a MacBook Air M4 CPU (2 cores), no GPU needed. Gated download.

Audio in

Text plus voice prompt

Audio out

PCM tensor at tts_model.sample_rate

Languages

en, fr, de, pt, it, es (model card)

Voices

Pre-made voices and cloning from a local wav

Latency

README: ~200 ms to first audio chunk; ~6x real time on M4 CPU with 2 cores.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

CPU (2 cores); no GPU needed

Licence

CC-BY-4.0 (model card)

Features

  • CPU only
  • audio streaming
  • voice cloning
  • built-in 'serve' command

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. pip install pocket-tts (on Linux use the CPU PyTorch index).
  2. Accept the gated model on Hugging Face.
  3. Use the Python API or 'pocket-tts serve'.

Endpoint

Local

Authentication

Hugging Face token

Quick start python

# pip install pocket-tts
from pocket_tts import TTSModel
import scipy.io.wavfile

tts_model = TTSModel.load_model()
voice_state = tts_model.get_state_for_audio_prompt("alba")  # or a local wav to clone
audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.")
scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Gated download

Access to kyutai/pocket-tts requires approval on Hugging Face.

Cloning consent

Voice cloning from any wav needs speaker consent.

Load voices once

load_model and get_state_for_audio_prompt are slow; cache them in memory.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Gated model access

Models and products

NameStatusNotes
kyutai/pocket-ttsGated on Hugging Face; updated 2026-10-01Also pocket-tts-without-voice-cloning.

Docs and sources

Docs

Sources used

Not fully verified

Code licence.

Similar open models APIs

Spotted a wrong price or a dead link?