GA OpenAI

OpenAI Text-to-Speech (gpt-4o-mini-tts)

Steerable TTS via /v1/audio/speech: you pass natural-language 'instructions' for tone and style. Output streams over chunked HTTP; there is no text-input WebSocket for TTS (use the Realtime API for that).

Est. per minute$0.013 - 0.027
1 high-severity warning

Overview

Best for: Teams already on OpenAI that want steerable, prompt-styled speech for non-interactive or sentence-chunked use.

At a glance

$/1M chars$15
Free tierNo
CloningYes
Text stream inNo
EmotionYes
8 kHz phoneNo
WebRTCNo
WebSocketNo
gRPCNo
Self-hostNo
Open weightsNo

$15/1M is tts-1 (tts-1-hd $30/1M). gpt-4o-mini-tts is token-billed ($12 per 1M audio tokens). HTTP chunked and SSE only; full text needed up front. tts-1 has 9 voices, gpt-4o-mini-tts has more. Custom voices limited to eligible customers with consent. No mulaw output (PCM is 24 kHz).

Audio in

Text plus optional free-text instructions

Audio out

mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless)

Languages

Follows Whisper language support; voices optimised for English

Voices

13 built-in (alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, cedar); custom voices for eligible customers with a consent recording

Latency

No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes.

Regions

Not specified

Compliance

OpenAI platform terms; disclosure of AI voice required by usage policy.

Features

  • output streaming
  • style instructions (tone, accent, whispering, speed)
  • custom voices (eligible customers, consent required)
  • OpenAI-compatible endpoint copied by many vendors

Pricing

WhatPriceUnitNotes
gpt-4o-mini-tts text input$0.60per 1M text tokens
gpt-4o-mini-tts audio output$12.00per 1M audio tokens
tts-1$15.00per 1M characters
tts-1-hd$30.00per 1M characters
How the per-minute estimate was worked out

900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now).

Free tier: None specific to TTS

Source: developers.openai.com

Setup

  1. Create an API key.
  2. Call POST https://api.openai.com/v1/audio/speech with model, voice, input and optional instructions.
  3. Use response_format pcm or wav and iterate the streamed body.
  4. Chunk long LLM output into sentences yourself (no input streaming).

Endpoint

https://api.openai.com/v1/audio/speech

Authentication

Authorization: Bearer <OPENAI_API_KEY>

Quick start python

# pip install openai
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY

with client.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",
    voice="marin",
    input="Hello! Your order has shipped and should arrive on Tuesday.",
    instructions="Warm, upbeat customer-support tone.",
    response_format="pcm",  # 24 kHz 16-bit mono, no header
) as resp:
    with open("out_24k_s16le.pcm", "wb") as f:
        for chunk in resp.iter_bytes(4096):
            f.write(chunk)  # or feed to a player as it arrives

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No text-input streaming

/v1/audio/speech needs the full input up front. For LLM token streams you must split into sentences and fire one request per sentence (watch rate limits), or use the Realtime API instead.

Instruction-following varies by snapshot

Developers report the 2025-12-15 snapshot follows style instructions (e.g. whispering) less consistently than 2025-03-20. Pin a dated snapshot and test your prompts.

Token billing makes cost harder to predict

gpt-4o-mini-tts bills audio output tokens, so slower speech or long pauses cost more than the character count suggests. Measure on your own content.

Disclosure is required

OpenAI usage policies require telling end users the voice is AI-generated. Custom voices need a recorded consent statement from the speaker.

Legacy models have fewer voices

tts-1 and tts-1-hd support only 9 voices; marin and cedar are gpt-4o-mini-tts only.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • gpt-4o-mini-tts max 2,000 input tokens per request
  • Rate limits by tier: Build 2,000 RPM / 150K TPM; Launch 10,000 RPM / 2M TPM; Grow 10,000 RPM / 8M TPM

Models and products

NameStatusNotes
gpt-4o-mini-tts (alias -> gpt-4o-mini-tts-2025-12-15)GADefault snapshot; supports instructions and custom voices. 2,000 max input tokens.
gpt-4o-mini-tts-2025-03-20Older snapshotSome developers report it follows style instructions better; community reports say it is being deprecated (date not confirmed).
tts-1GA (legacy)Lower latency, lower quality; 9 voices.
tts-1-hdGA (legacy)Higher quality; 9 voices.

Docs and sources

Docs

Sources used

Not fully verified

Per-minute cost for gpt-4o-mini-tts; deprecation date of the 2025-03-20 snapshot; SSE stream_format option.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?