GA Resemble AI

Chatterbox (Turbo, Nano, Multilingual V3)

MIT-licensed open TTS family with zero-shot cloning: Turbo (350M, English, one-step decoder, paralinguistic tags), Nano (110M, CPU) and Multilingual V3 (500M, 23+ languages). Outputs are watermarked.

Est. per minuten/a
1 high-severity warning

Overview

Best for: Commercial-friendly self-hosted cloning; English agents with Turbo; CPU/edge with Nano.

At a glance

TypeText-to-speech
Params B0.35
CPU okYes
LicenceMIT
CommercialYes
StreamingNo
Latency ms103
Languages23

103 ms is time to first packet for the Flash variant per its card. 23+ languages on Multilingual V3 (Turbo and Nano English only). Params: Turbo 350M (Nano 110M runs on CPU, Multilingual 500M). No first-party streaming in the library. Outputs watermarked.

Audio in

Text plus optional reference audio prompt

Audio out

Waveform at model sample rate (model.sr)

Languages

Turbo/Nano: English; Multilingual V3: 23+

Voices

Zero-shot cloning from a reference clip

Latency

Hosted Resemble service claims sub-200 ms; self-hosted depends on GPU. No official streaming API in the library.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

Turbo: modest GPU; Nano: CPU (8 cores for 3x real time); Multilingual: GPU.

Licence

MIT (Turbo, Nano, Flash, base); Dramabox model is 'other'

Features

  • zero-shot voice cloning
  • paralinguistic tags (Turbo/Nano)
  • PerTh watermark
  • CPU model (Nano)

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. pip install chatterbox-tts.
  2. Load ChatterboxTurboTTS on cuda (or Nano on CPU).
  3. Generate per sentence with an optional reference clip.

Endpoint

Local

Authentication

None

Quick start python

# pip install chatterbox-tts
import torchaudio as ta
from chatterbox.tts_turbo import ChatterboxTurboTTS

model = ChatterboxTurboTTS.from_pretrained(device="cuda")
wav = model.generate(
    "Hi there [chuckle], thanks for calling back.",
    audio_prompt_path="reference_voice.wav",  # zero-shot clone (get consent)
)
ta.save("out.wav", wav, model.sr)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Zero-shot cloning needs consent

Anyone's voice can be cloned from a short clip. Get written consent and keep records; impersonation laws apply regardless of the open licence.

Built-in watermark

Every output carries Resemble's PerTh imperceptible watermark. Usually harmless, but you cannot remove it within the licence spirit.

No first-party streaming

The library returns full utterances; for agents, chunk by sentence or use a community streaming fork.

Turbo and Nano are English only

Use Multilingual V3 for other languages, at higher compute.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Library generates whole utterances; streaming needs community forks or your own chunking

Models and products

NameStatusNotes
ResembleAI/chatterbox-turboReleased 2025-12350M, English, [laugh]/[cough] tags.
ResembleAI/chatterbox-nanoReleased 2026-04110M, 3x real time on 8 CPU cores (README).
ResembleAI/chatterbox-flashReleased 2026-05Block-diffusion variant; model card reports time-to-first-packet from 103 ms.
Chatterbox-Multilingual V3 + single-language packsReleased 2026-06500M, 23+ languages.

Docs and sources

Docs

Sources used

Not fully verified

Self-hosted latency numbers.

Similar open models APIs

Spotted a wrong price or a dead link?