Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Azure AI Speech neural and HD voices
Microsoft
OpenAI Text-to-Speech (gpt-4o-mini-tts)
OpenAI
Fish Audio API (S2.1 Pro)
Fish Audio
CategoryText-to-speechText-to-speechText-to-speech
StatusGAGAGA
Est. per minute$0.013 - 0.02$0.013 - 0.027$0.013 - 0.041
How that was worked out900 chars/min; Neural $15/1M vs Neural HD $22/1M at pay-as-you-go900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now).900 chars/min. English ASCII = 1 byte/char ($0.0135/min); CJK/Arabic/Hindi are ~3 bytes/char (~$0.04/min). Free model excluded.
Pricing modelper-characterper-tokenper-character
Free tierF0: 0.5M neural characters per monthNone specific to TTSs2.1-pro-free model at $0 under fair use (time-limited per third-party sources)
Connects byWebSocket, HTTP chunkedHTTP chunked, SSEWebSocket, HTTP chunked
Audio inText or SSML (text streaming mode does not support SSML)Text plus optional free-text instructionsText (MessagePack frames on WebSocket)
Audio outopus, mp3, pcm, truesilk at 8/16/24/48 kHz; raw PCM formats such as Raw24Khz16BitMonoPcmmp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless)mp3 (default, 64/128/192 kbps), wav, pcm, opus; 44.1 kHz for most formats, 48 kHz for opus
LanguagesMany locales; see language-support page (count not re-verified)Follows Whisper language support; voices optimised for English83 per third-party coverage of S2.1 Pro (not on the docs pages checked)
Latency (vendor claim)Microsoft comparison table: HD and standard neural voices < 300 ms; Azure OpenAI voices > 500 ms.No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes.Vendor claim (third-party reported): ~90 ms to first audio for S2.1 Pro; Vapi measured 141 ms median including network.
Key limits
  • Text streaming: SDK only (C#, C++, Python), WebSocket v2 endpoint required
  • HD voices: real-time only, subset of SSML, cloud only
  • Concurrency per resource defaults; see quotas page
  • gpt-4o-mini-tts max 2,000 input tokens per request
  • Rate limits by tier: Build 2,000 RPM / 150K TPM; Launch 10,000 RPM / 2M TPM; Grow 10,000 RPM / 8M TPM
  • Concurrent requests: Starter (<$100 paid) 5, Elevated ($100+) 15, High Volume ($1,000+) 50
  • 429 without Retry-After; use exponential backoff
  • chunk_length 100-300
High-severity warnings
  • Text streaming needs the SDK and the v2 endpoint
  • No text-input streaming
  • Billed per UTF-8 byte, not character
ComplianceAzure compliance programs; containers and disconnected options for non-HD voices. Specific certifications not re-verified here.OpenAI platform terms; disclosure of AI voice required by usage policy.Not verified.
Self-hostableYesNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/1M chars$15$15$15
Free tierYesNoYes
Free credit $---
Free tier commercial---
Voices500--
CloningYesYesYes
Instant clone---
Text stream inYesNoYes
TimestampsYes--
EmotionYesYes-
SSMLYes--
8 kHz phone-NoNo
Latency ms300-90
Languages--83
Max session min---
Concurrency--5
WebRTCNoNoNo
WebSocketYesNoYes
gRPCNoNoNo
HIPAAYes--
SOC 2Yes--
EU dataYes--
Self-hostYesNoNo
Open weightsNoNoYes
High warnings111