Qwen3-TTS
Apache-2.0 open TTS (0.6B and 1.7B) released January 2026 with dual-track streaming that can emit audio after a single input character, 10 languages, instruction control, voice design and cloning.
Overview
Best for: Commercial-friendly self-hosted multilingual (especially Chinese/English) TTS with instructions.
At a glance
97 ms is the vendor end-to-end figure; the simple Python API returns whole utterances. Also 0.6B variants.
Text plus optional instruction
Waveform (sample rate returned by generate)
10: zh, en, ja, ko, de, fr, ru, pt, es, it
Preset speakers, voice design, voice cloning
README: end-to-end synthesis latency as low as 97 ms; first packet after one character.
Wherever you deploy it
Your own deployment; no vendor data processing
NVIDIA GPU (bf16, flash-attention recommended)
Apache-2.0
Features
- streaming generation
- instruction-controlled emotion and pace
- voice design
- voice cloning
- vLLM-Omni support (offline at launch)
Pricing
| What | Price | Unit |
|---|---|---|
| Weights | $0 |
Self-hosted: cost is your GPU/CPU time, not per character
Free tier: Open weights
Source: huggingface.co
Setup
- pip install -U qwen-tts (flash-attn recommended).
- Load a model from Hugging Face.
- Call generate_custom_voice / voice design / clone functions.
Endpoint
Local
Authentication
None
Quick start python
# pip install -U qwen-tts soundfile
import torch, soundfile as sf
from qwen_tts import Qwen3TTSModel
model = Qwen3TTSModel.from_pretrained("Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
device_map="cuda:0", dtype=torch.bfloat16, attn_implementation="flash_attention_2")
wavs, sr = model.generate_custom_voice(
text="Your appointment is confirmed for Tuesday at ten.",
language="English", speaker="Vivian", instruct="Calm and friendly.")
sf.write("out.wav", wavs[0], sr)
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Streaming needs the right serving path
The 97 ms figure is from the vendor; the simple Python API returns whole utterances. Real streaming depends on the serving stack you choose.
Cloning consent
Voice cloning of real people requires consent.
flash-attn build pain
flash-attn often needs compiling; set MAX_JOBS to avoid running out of RAM.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- vLLM-Omni online serving was 'coming later' at launch
Models and products
| Name | Status |
|---|---|
| Qwen/Qwen3-TTS-12Hz-1.7B-Base / -CustomVoice / 0.6B variants | Released 2026-01-21 |
Docs and sources
Docs
Sources used
Independent latency.