GA Alibaba Qwen

Qwen3-TTS

Apache-2.0 open TTS (0.6B and 1.7B) released January 2026 with dual-track streaming that can emit audio after a single input character, 10 languages, instruction control, voice design and cloning.

Est. per minuten/a

Overview

Best for: Commercial-friendly self-hosted multilingual (especially Chinese/English) TTS with instructions.

At a glance

TypeText-to-speech
Params B1.7
LicenceApache-2.0
CommercialYes
StreamingYes
Latency ms97
Languages10

97 ms is the vendor end-to-end figure; the simple Python API returns whole utterances. Also 0.6B variants.

Audio in

Text plus optional instruction

Audio out

Waveform (sample rate returned by generate)

Languages

10: zh, en, ja, ko, de, fr, ru, pt, es, it

Voices

Preset speakers, voice design, voice cloning

Latency

README: end-to-end synthesis latency as low as 97 ms; first packet after one character.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

NVIDIA GPU (bf16, flash-attention recommended)

Licence

Apache-2.0

Features

  • streaming generation
  • instruction-controlled emotion and pace
  • voice design
  • voice cloning
  • vLLM-Omni support (offline at launch)

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. pip install -U qwen-tts (flash-attn recommended).
  2. Load a model from Hugging Face.
  3. Call generate_custom_voice / voice design / clone functions.

Endpoint

Local

Authentication

None

Quick start python

# pip install -U qwen-tts soundfile
import torch, soundfile as sf
from qwen_tts import Qwen3TTSModel

model = Qwen3TTSModel.from_pretrained("Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    device_map="cuda:0", dtype=torch.bfloat16, attn_implementation="flash_attention_2")
wavs, sr = model.generate_custom_voice(
    text="Your appointment is confirmed for Tuesday at ten.",
    language="English", speaker="Vivian", instruct="Calm and friendly.")
sf.write("out.wav", wavs[0], sr)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Streaming needs the right serving path

The 97 ms figure is from the vendor; the simple Python API returns whole utterances. Real streaming depends on the serving stack you choose.

Cloning consent

Voice cloning of real people requires consent.

flash-attn build pain

flash-attn often needs compiling; set MAX_JOBS to avoid running out of RAM.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • vLLM-Omni online serving was 'coming later' at launch

Models and products

NameStatusNotes
Qwen/Qwen3-TTS-12Hz-1.7B-Base / -CustomVoice / 0.6B variantsReleased 2026-01-21

Docs and sources

Docs

Sources used

Not fully verified

Independent latency.

Similar open models APIs

Spotted a wrong price or a dead link?