Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| OpenAI Text-to-Speech (gpt-4o-mini-tts) OpenAI | Azure AI Speech neural and HD voices Microsoft | Fish Audio API (S2.1 Pro) Fish Audio | |
|---|---|---|---|
| Category | Text-to-speech | Text-to-speech | Text-to-speech |
| Status | GA | GA | GA |
| Est. per minute | $0.013 - 0.027 | $0.013 - 0.02 | $0.013 - 0.041 |
| How that was worked out | 900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now). | 900 chars/min; Neural $15/1M vs Neural HD $22/1M at pay-as-you-go | 900 chars/min. English ASCII = 1 byte/char ($0.0135/min); CJK/Arabic/Hindi are ~3 bytes/char (~$0.04/min). Free model excluded. |
| Pricing model | per-token | per-character | per-character |
| Free tier | None specific to TTS | F0: 0.5M neural characters per month | s2.1-pro-free model at $0 under fair use (time-limited per third-party sources) |
| Connects by | HTTP chunked, SSE | WebSocket, HTTP chunked | WebSocket, HTTP chunked |
| Audio in | Text plus optional free-text instructions | Text or SSML (text streaming mode does not support SSML) | Text (MessagePack frames on WebSocket) |
| Audio out | mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless) | opus, mp3, pcm, truesilk at 8/16/24/48 kHz; raw PCM formats such as Raw24Khz16BitMonoPcm | mp3 (default, 64/128/192 kbps), wav, pcm, opus; 44.1 kHz for most formats, 48 kHz for opus |
| Languages | Follows Whisper language support; voices optimised for English | Many locales; see language-support page (count not re-verified) | 83 per third-party coverage of S2.1 Pro (not on the docs pages checked) |
| Latency (vendor claim) | No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes. | Microsoft comparison table: HD and standard neural voices < 300 ms; Azure OpenAI voices > 500 ms. | Vendor claim (third-party reported): ~90 ms to first audio for S2.1 Pro; Vapi measured 141 ms median including network. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | OpenAI platform terms; disclosure of AI voice required by usage policy. | Azure compliance programs; containers and disconnected options for non-HD voices. Specific certifications not re-verified here. | Not verified. |
| Self-hostable | No | Yes | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| $/1M chars | $15 | $15 | $15 |
| Free tier | No | Yes | Yes |
| Free credit $ | - | - | - |
| Free tier commercial | - | - | - |
| Voices | - | 500 | - |
| Cloning | Yes | Yes | Yes |
| Instant clone | - | - | - |
| Text stream in | No | Yes | Yes |
| Timestamps | - | Yes | - |
| Emotion | Yes | Yes | - |
| SSML | - | Yes | - |
| 8 kHz phone | No | - | No |
| Latency ms | - | 300 | 90 |
| Languages | - | - | 83 |
| Max session min | - | - | - |
| Concurrency | - | - | 5 |
| WebRTC | No | No | No |
| WebSocket | No | Yes | Yes |
| gRPC | No | No | No |
| HIPAA | - | Yes | - |
| SOC 2 | - | Yes | - |
| EU data | - | Yes | - |
| Self-host | No | Yes | No |
| Open weights | No | No | Yes |
| High warnings | 1 | 1 | 1 |