Kokoro-82M
Tiny 82M-parameter open TTS with Apache-2.0 weights; fast enough for CPU and very fast on GPU. The default cheap self-hosted voice for many hobby and production stacks.
Overview
Best for: Cheapest decent-quality self-hosted English voice; edge and CPU deployments.
At a glance
Sentence-segment output; no text-input streaming. 54 voices, no cloning.
Text (phonemized via misaki/espeak)
24 kHz float audio from the Python library; community servers expose wav/mp3/pcm
8 (American/British English, plus others listed in VOICES.md)
54 preset voices; no voice cloning
No official TTFB figure; generates per sentence segment.
Wherever you deploy it
Your own deployment; no vendor data processing
Runs on CPU; any small GPU makes it much faster than real time. 82M parameters.
Apache-2.0 (weights and code)
Features
- sentence-segment generator output
- OpenAI-compatible community servers (e.g. Kokoro-FastAPI)
- CPU capable
Pricing
| What | Price | Unit |
|---|---|---|
| Weights | $0 |
Self-hosted: cost is your GPU/CPU time, not per character
Free tier: Open weights
Source: huggingface.co
Setup
- pip install kokoro soundfile (and install espeak-ng for fallback/non-English).
- Create a KPipeline with a lang_code.
- Iterate the generator; each yield is one audio segment you can play immediately.
Endpoint
Local (library) or your own HTTP server
Authentication
None (add your own)
Quick start python
# pip install kokoro soundfile (+ apt install espeak-ng)
from kokoro import KPipeline
import soundfile as sf
pipeline = KPipeline(lang_code="a") # 'a' = American English
text = "Hello there. Kokoro yields audio sentence by sentence, so playback can start early."
for i, (graphemes, phonemes, audio) in enumerate(pipeline(text, voice="af_heart")):
sf.write(f"part_{i}.wav", audio, 24000) # stream each part to the player instead
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
No cloning, fixed voices
Only the shipped voices; you cannot clone a brand voice.
Text-in streaming is your job
The pipeline splits complete text into segments. For LLM streams, buffer tokens to sentence boundaries and call it per sentence.
espeak-ng dependency
Out-of-dictionary English and several non-English languages need espeak-ng installed; missing it causes silent mispronunciations or errors.
Community servers vary
OpenAI-compatible wrappers are third-party projects; check their licences and maintenance.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- No native text-input streaming: you feed sentences
- Quality drops on unusual words (espeak fallback)
Models and products
| Name | Status |
|---|---|
| hexgrad/Kokoro-82M v1.0 | Released 2025-01-27 |
Docs and sources
Docs
Sources used
Latency on specific hardware.