GA Speechify

SpeechifyAI (Simba 3.2 / 3.0)

Developer API from the Speechify reader company. Cheap per-character pricing ($6-10 per 1M) and a streaming HTTP endpoint; Simba 3.2 is English-only and streaming-native. No text-input WebSocket documented.

Est. per minute$0.0054 - 0.009

Overview

Best for: Low-cost English narration and agents where HTTP streaming per sentence is acceptable.

At a glance

$/1M chars$5.26
Free tierYes
Free tier commercialYes
CloningYes
Text stream inNo
SSMLYes
Latency ms56
Languages6
WebRTCNo
WebSocketNo
gRPCNo
Self-hostNo
Open weightsNo

56 ms p50 vendor claim (independent medians 106-123 ms). Price is the Starter plan effective rate ($10 for 1.9M chars); overage $10/1M, Scale $6/1M. simba-3.2 is English only; simba-3.0 covers 6 languages. HTTP output streaming only. Free 500K chars/month.

Audio in

Text and SSML (pauses, speaking rate)

Audio out

5 formats including mp3 (per marketing page)

Languages

English (simba-3.2) plus 5 more on simba-3.0

Voices

Catalog voices (e.g. geffen_32); cloning on paid plans, cloned voices on simba-3.2 need manual approval

Latency

Vendor: 56 ms p50 first byte; independent: 106 ms (Coval) and 123 ms (Voice Arena) median time to first audio.

Regions

Production US East referenced for latency

Compliance

Not verified.

Features

  • streaming HTTP output
  • SSML
  • voice cloning (paid)
  • batch synthesis (Scale)
  • docs MCP server

Pricing

WhatPriceUnitNotes
Free$0monthly500K characters/month, no card; hard stop at balance
Starter$10/monthmonthly1.9M chars included, then $10 per 1M
Pro$99/monthmonthly13.5M chars included, then $8 per 1M
Scale$499/monthmonthly78M chars included, then $6 per 1M
How the per-minute estimate was worked out

900 chars/min at $6 (Scale) to $10 (Starter) per 1M

Free tier: 500K characters/month; commercial use on every plan (vendor)

Source: speechify.ai

Setup

  1. Sign up at platform.speechify.ai and create an API key.
  2. POST https://api.speechify.ai/v1/audio/stream with input, voice_id and model.
  3. Stream the response body to a player.

Endpoint

https://api.speechify.ai/v1/audio/stream

Authentication

Authorization: Bearer <SPEECHIFY_API_KEY>

Quick start python

# pip install requests
import os, requests

r = requests.post(
    "https://api.speechify.ai/v1/audio/stream",
    headers={"Authorization": f"Bearer {os.environ['SPEECHIFY_API_KEY']}",
             "Content-Type": "application/json", "Accept": "audio/mpeg"},
    json={"input": "Streaming speech from the Speechify API.",
          "voice_id": "geffen_32", "model": "simba-3.2"},
    stream=True, timeout=60)
r.raise_for_status()
with open("out.mp3", "wb") as f:
    for chunk in r.iter_content(4096):
        f.write(chunk)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No text-input streaming

Only HTTP output streaming is documented; you must sentence-chunk LLM output and issue one request per chunk.

simba-3.2 is English only

Multilingual needs simba-3.0, which is also the silent default when you omit model.

Old model IDs fail

simba-english and simba-multilingual return 400 model_retired on new workspaces; update older sample code.

Do not confuse with the consumer app

speechify.com is the reader app with separate billing; the API lives at speechify.ai / api.speechify.ai.

Accept header for MP3 is assumed

The snippet sets Accept: audio/mpeg; check the API reference for the exact format selector.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • /v1/audio/stream up to 20,000 characters; /v1/audio/speech up to 2,000
  • Free tier cannot top up

Models and products

NameStatusNotes
simba-3.2GA (recommended for English)Vendor: 56 ms p50 first byte (US East, 2026-09-15); Coval 106 ms and Voice Arena 123 ms medians including network (2026-09-24).
simba-3.0GA (default when model omitted)English, German, Spanish, French, Italian, Portuguese.
simba-english / simba-multilingualRetiredNew workspaces get 400 model_retired.

Docs and sources

Docs

Sources used

Not fully verified

Exact audio format parameter name and list; concurrency limits.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?