Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

OpenAI Text-to-Speech (gpt-4o-mini-tts)
OpenAI
Azure AI Speech neural and HD voices
Microsoft
Fish Audio API (S2.1 Pro)
Fish Audio
CategoryText-to-speechText-to-speechText-to-speech
StatusGAGAGA
Est. per minute$0.013 - 0.027$0.013 - 0.02$0.013 - 0.041
How that was worked out900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now).900 chars/min; Neural $15/1M vs Neural HD $22/1M at pay-as-you-go900 chars/min. English ASCII = 1 byte/char ($0.0135/min); CJK/Arabic/Hindi are ~3 bytes/char (~$0.04/min). Free model excluded.
Pricing modelper-tokenper-characterper-character
Free tierNone specific to TTSF0: 0.5M neural characters per months2.1-pro-free model at $0 under fair use (time-limited per third-party sources)
Connects byHTTP chunked, SSEWebSocket, HTTP chunkedWebSocket, HTTP chunked
Audio inText plus optional free-text instructionsText or SSML (text streaming mode does not support SSML)Text (MessagePack frames on WebSocket)
Audio outmp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless)opus, mp3, pcm, truesilk at 8/16/24/48 kHz; raw PCM formats such as Raw24Khz16BitMonoPcmmp3 (default, 64/128/192 kbps), wav, pcm, opus; 44.1 kHz for most formats, 48 kHz for opus
LanguagesFollows Whisper language support; voices optimised for EnglishMany locales; see language-support page (count not re-verified)83 per third-party coverage of S2.1 Pro (not on the docs pages checked)
Latency (vendor claim)No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes.Microsoft comparison table: HD and standard neural voices < 300 ms; Azure OpenAI voices > 500 ms.Vendor claim (third-party reported): ~90 ms to first audio for S2.1 Pro; Vapi measured 141 ms median including network.
Key limits
  • gpt-4o-mini-tts max 2,000 input tokens per request
  • Rate limits by tier: Build 2,000 RPM / 150K TPM; Launch 10,000 RPM / 2M TPM; Grow 10,000 RPM / 8M TPM
  • Text streaming: SDK only (C#, C++, Python), WebSocket v2 endpoint required
  • HD voices: real-time only, subset of SSML, cloud only
  • Concurrency per resource defaults; see quotas page
  • Concurrent requests: Starter (<$100 paid) 5, Elevated ($100+) 15, High Volume ($1,000+) 50
  • 429 without Retry-After; use exponential backoff
  • chunk_length 100-300
High-severity warnings
  • No text-input streaming
  • Text streaming needs the SDK and the v2 endpoint
  • Billed per UTF-8 byte, not character
ComplianceOpenAI platform terms; disclosure of AI voice required by usage policy.Azure compliance programs; containers and disconnected options for non-HD voices. Specific certifications not re-verified here.Not verified.
Self-hostableNoYesNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/1M chars$15$15$15
Free tierNoYesYes
Free credit $---
Free tier commercial---
Voices-500-
CloningYesYesYes
Instant clone---
Text stream inNoYesYes
Timestamps-Yes-
EmotionYesYes-
SSML-Yes-
8 kHz phoneNo-No
Latency ms-30090
Languages--83
Max session min---
Concurrency--5
WebRTCNoNoNo
WebSocketNoYesYes
gRPCNoNoNo
HIPAA-Yes-
SOC 2-Yes-
EU data-Yes-
Self-hostNoYesNo
Open weightsNoNoYes
High warnings111