F5-TTS
Popular flow-matching zero-shot cloning model. Code is MIT but the pretrained weights are CC-BY-NC because of the Emilia training data.
Overview
Best for: Research and non-commercial cloning experiments.
At a glance
Code MIT, pretrained weights CC-BY-NC-4.0. English and Chinese base.
Text plus reference audio and its transcript
24 kHz waveform
English and Chinese base; community finetunes for others
Zero-shot cloning
Non-autoregressive; no first-party streaming server.
Wherever you deploy it
Your own deployment; no vendor data processing
GPU
Code MIT; pretrained weights CC-BY-NC-4.0
Features
- zero-shot cloning
- Docker image
- CLI and Gradio app
Pricing
| What | Price | Unit |
|---|---|---|
| Weights | $0 |
Self-hosted: cost is your GPU/CPU time, not per character
Free tier: Open weights
Source: huggingface.co
Setup
- pip install f5-tts (GPU).
- Run f5-tts_infer-cli --model F5TTS_v1_Base with a reference clip.
Endpoint
Local
Authentication
None
Warnings
Weights are non-commercial
The official checkpoints are CC-BY-NC due to Emilia training data. Commercial products need your own retrained weights.
Not a streaming engine
Generates whole utterances; poor fit for live agents.
Cloning consent
Zero-shot cloning of real people needs consent.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Non-commercial weights
Models and products
| Name | Status |
|---|---|
| F5TTS_v1_Base | Released 2025-03-12 |