GA Fish Audio

Fish Audio API (S2.1 Pro)

Multilingual (83 languages per third parties) cloning-first TTS with a MessagePack WebSocket for streamed text. Billed per UTF-8 byte, with a free s2.1-pro-free model under fair use.

Est. per minute$0.013 - 0.041
1 high-severity warning

Overview

Best for: Multilingual cloning and character voices at a low flat price.

At a glance

$/1M chars$15
Free tierYes
CloningYes
Text stream inYes
8 kHz phoneNo
Latency ms90
Languages83
Concurrency5
WebRTCNo
WebSocketYes
gRPCNo
Self-hostNo
Open weightsYes

$15 per 1M UTF-8 bytes, so non-Latin scripts cost about 3x per character. Latency (~90 ms) and 83 languages come from third-party coverage. Free model s2.1-pro-free has no SLA, is time-limited and may retain data. Open weights (S2 Pro) are non-commercial. Concurrency 5 until $100 paid.

Audio in

Text (MessagePack frames on WebSocket)

Audio out

mp3 (default, 64/128/192 kbps), wav, pcm, opus; 44.1 kHz for most formats, 48 kHz for opus

Languages

83 per third-party coverage of S2.1 Pro (not on the docs pages checked)

Voices

Very large community voice library; instant cloning via references

Latency

Vendor claim (third-party reported): ~90 ms to first audio for S2.1 Pro; Vapi measured 141 ms median including network.

Regions

Not specified

Compliance

Not verified.

Features

  • input streaming (MessagePack WebSocket)
  • latency modes low/balanced/normal
  • voice cloning from reference audio
  • multi-speaker dialogue
  • prosody speed/volume

Pricing

WhatPriceUnitNotes
s2.1-pro / s2-pro / s1$15.00per 1M UTF-8 bytesNo subscription or minimum for API
s2.1-pro-free$0.00per 1M UTF-8 bytesFair use, no SLA
voice-design-1$0.01per successful request
How the per-minute estimate was worked out

900 chars/min. English ASCII = 1 byte/char ($0.0135/min); CJK/Arabic/Hindi are ~3 bytes/char (~$0.04/min). Free model excluded.

Free tier: s2.1-pro-free model at $0 under fair use (time-limited per third-party sources)

Source: docs.fish.audio

Setup

  1. Create an API key and prepay credit at fish.audio.
  2. Choose a voice reference_id (or upload references).
  3. Open wss://api.fish.audio/v1/tts/live with Authorization: Bearer and a model header.
  4. Send start, text..., flush, stop as MessagePack; collect 'audio' events until 'finish'.
  5. Or use the Python SDK's stream_websocket with a text generator.

Endpoint

wss://api.fish.audio/v1/tts/live

Authentication

Authorization: Bearer <FISH_API_KEY>; header model: s2.1-pro

Quick start python

# Fish Audio Python SDK (import name 'fishaudio'; check PyPI for the package name)
from fishaudio import FishAudio

client = FishAudio()  # reads FISH_API_KEY

def llm_tokens():
    for token in ["The ", "first ", "move ", "sets ", "everything ", "in ", "motion."]:
        yield token

with open("out.mp3", "wb") as f:
    for chunk in client.tts.stream_websocket(llm_tokens(), reference_id="YOUR_VOICE_ID"):
        f.write(chunk)  # play or forward as it arrives

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Billed per UTF-8 byte, not character

Non-Latin scripts cost ~3x per character (Chinese, Japanese, Korean, Arabic, Hindi are 3-4 bytes each). Compare vendors on bytes for non-English workloads.

Free model is temporary and may train on your data

s2.1-pro-free has no SLA, is reported to run only through 2026-11-30, and Fish's blog says requests may be retained for model improvement. Do not send sensitive text.

s1 retires 2026-12-31

After that date s1 requests are served by s2.1-pro; voices may sound different.

Low default concurrency

Only 5 concurrent requests until you have paid $100; 429s carry no Retry-After header.

Cloning community voices

The public voice library contains user-uploaded clones; confirm you have rights to any voice you use commercially.

Open weights are non-commercial

The S2 Pro weights on Hugging Face use the Fish Audio Research License, not a commercial licence; self-hosting for a product needs a separate agreement.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • Concurrent requests: Starter (<$100 paid) 5, Elevated ($100+) 15, High Volume ($1,000+) 50
  • 429 without Retry-After; use exponential backoff
  • chunk_length 100-300

Models and products

NameStatusNotes
s2.1-proGA (default if model header omitted)$15/M UTF-8 bytes.
s2.1-pro-freeFree, fair use, no SLAThird-party sources say available through 2026-11-30; requests may be retained for training.
s2-proGAMulti-speaker dialogue supported.
s1Deprecated, retiring 2026-12-31After that, s1 requests are served and billed as s2.1-pro.

Docs and sources

Docs

Sources used

Not fully verified

Language count (83) and ~90 ms latency come from third-party coverage; end date of the free model; Python package name.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?