GA hexgrad (community)

Kokoro-82M

Tiny 82M-parameter open TTS with Apache-2.0 weights; fast enough for CPU and very fast on GPU. The default cheap self-hosted voice for many hobby and production stacks.

Est. per minuten/a

Overview

Best for: Cheapest decent-quality self-hosted English voice; edge and CPU deployments.

At a glance

TypeText-to-speech
Params B0.082
CPU okYes
LicenceApache-2.0
CommercialYes
StreamingNo
Languages8

Sentence-segment output; no text-input streaming. 54 voices, no cloning.

Audio in

Text (phonemized via misaki/espeak)

Audio out

24 kHz float audio from the Python library; community servers expose wav/mp3/pcm

Languages

8 (American/British English, plus others listed in VOICES.md)

Voices

54 preset voices; no voice cloning

Latency

No official TTFB figure; generates per sentence segment.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

Runs on CPU; any small GPU makes it much faster than real time. 82M parameters.

Licence

Apache-2.0 (weights and code)

Features

  • sentence-segment generator output
  • OpenAI-compatible community servers (e.g. Kokoro-FastAPI)
  • CPU capable

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. pip install kokoro soundfile (and install espeak-ng for fallback/non-English).
  2. Create a KPipeline with a lang_code.
  3. Iterate the generator; each yield is one audio segment you can play immediately.

Endpoint

Local (library) or your own HTTP server

Authentication

None (add your own)

Quick start python

# pip install kokoro soundfile   (+ apt install espeak-ng)
from kokoro import KPipeline
import soundfile as sf

pipeline = KPipeline(lang_code="a")  # 'a' = American English
text = "Hello there. Kokoro yields audio sentence by sentence, so playback can start early."
for i, (graphemes, phonemes, audio) in enumerate(pipeline(text, voice="af_heart")):
    sf.write(f"part_{i}.wav", audio, 24000)  # stream each part to the player instead

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No cloning, fixed voices

Only the shipped voices; you cannot clone a brand voice.

Text-in streaming is your job

The pipeline splits complete text into segments. For LLM streams, buffer tokens to sentence boundaries and call it per sentence.

espeak-ng dependency

Out-of-dictionary English and several non-English languages need espeak-ng installed; missing it causes silent mispronunciations or errors.

Community servers vary

OpenAI-compatible wrappers are third-party projects; check their licences and maintenance.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • No native text-input streaming: you feed sentences
  • Quality drops on unusual words (espeak fallback)

Models and products

NameStatusNotes
hexgrad/Kokoro-82M v1.0Released 2025-01-278 languages, 54 voices.

Docs and sources

Docs

Sources used

Not fully verified

Latency on specific hardware.

Similar open models APIs

Spotted a wrong price or a dead link?