GA SWivid (academic)

F5-TTS

Popular flow-matching zero-shot cloning model. Code is MIT but the pretrained weights are CC-BY-NC because of the Emilia training data.

Est. per minuten/a
1 high-severity warning

Overview

Best for: Research and non-commercial cloning experiments.

At a glance

TypeText-to-speech
LicenceCC-BY-NC-4.0
CommercialNo
StreamingNo
Languages2

Code MIT, pretrained weights CC-BY-NC-4.0. English and Chinese base.

Audio in

Text plus reference audio and its transcript

Audio out

24 kHz waveform

Languages

English and Chinese base; community finetunes for others

Voices

Zero-shot cloning

Latency

Non-autoregressive; no first-party streaming server.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

GPU

Licence

Code MIT; pretrained weights CC-BY-NC-4.0

Features

  • zero-shot cloning
  • Docker image
  • CLI and Gradio app

Pricing

WhatPriceUnitNotes
Weights$0Weights free for non-commercial use only
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. pip install f5-tts (GPU).
  2. Run f5-tts_infer-cli --model F5TTS_v1_Base with a reference clip.

Endpoint

Local

Authentication

None

Warnings

Weights are non-commercial

The official checkpoints are CC-BY-NC due to Emilia training data. Commercial products need your own retrained weights.

Not a streaming engine

Generates whole utterances; poor fit for live agents.

Cloning consent

Zero-shot cloning of real people needs consent.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Non-commercial weights

Models and products

NameStatusNotes
F5TTS_v1_BaseReleased 2025-03-12

Docs and sources

Docs

Sources used

Similar open models APIs

Spotted a wrong price or a dead link?