Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Fish Audio API (S2.1 Pro) Fish Audio | OpenAI Text-to-Speech (gpt-4o-mini-tts) OpenAI | Azure AI Speech neural and HD voices Microsoft | |
|---|---|---|---|
| Category | Text-to-speech | Text-to-speech | Text-to-speech |
| Status | GA | GA | GA |
| Est. per minute | $0.013 - 0.041 | $0.013 - 0.027 | $0.013 - 0.02 |
| How that was worked out | 900 chars/min. English ASCII = 1 byte/char ($0.0135/min); CJK/Arabic/Hindi are ~3 bytes/char (~$0.04/min). Free model excluded. | 900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now). | 900 chars/min; Neural $15/1M vs Neural HD $22/1M at pay-as-you-go |
| Pricing model | per-character | per-token | per-character |
| Free tier | s2.1-pro-free model at $0 under fair use (time-limited per third-party sources) | None specific to TTS | F0: 0.5M neural characters per month |
| Connects by | WebSocket, HTTP chunked | HTTP chunked, SSE | WebSocket, HTTP chunked |
| Audio in | Text (MessagePack frames on WebSocket) | Text plus optional free-text instructions | Text or SSML (text streaming mode does not support SSML) |
| Audio out | mp3 (default, 64/128/192 kbps), wav, pcm, opus; 44.1 kHz for most formats, 48 kHz for opus | mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless) | opus, mp3, pcm, truesilk at 8/16/24/48 kHz; raw PCM formats such as Raw24Khz16BitMonoPcm |
| Languages | 83 per third-party coverage of S2.1 Pro (not on the docs pages checked) | Follows Whisper language support; voices optimised for English | Many locales; see language-support page (count not re-verified) |
| Latency (vendor claim) | Vendor claim (third-party reported): ~90 ms to first audio for S2.1 Pro; Vapi measured 141 ms median including network. | No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes. | Microsoft comparison table: HD and standard neural voices < 300 ms; Azure OpenAI voices > 500 ms. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Not verified. | OpenAI platform terms; disclosure of AI voice required by usage policy. | Azure compliance programs; containers and disconnected options for non-HD voices. Specific certifications not re-verified here. |
| Self-hostable | No | No | Yes |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| $/1M chars | $15 | $15 | $15 |
| Free tier | Yes | No | Yes |
| Free credit $ | - | - | - |
| Free tier commercial | - | - | - |
| Voices | - | - | 500 |
| Cloning | Yes | Yes | Yes |
| Instant clone | - | - | - |
| Text stream in | Yes | No | Yes |
| Timestamps | - | - | Yes |
| Emotion | - | Yes | Yes |
| SSML | - | - | Yes |
| 8 kHz phone | No | No | - |
| Latency ms | 90 | - | 300 |
| Languages | 83 | - | - |
| Max session min | - | - | - |
| Concurrency | 5 | - | - |
| WebRTC | No | No | No |
| WebSocket | Yes | No | Yes |
| gRPC | No | No | No |
| HIPAA | - | - | Yes |
| SOC 2 | - | - | Yes |
| EU data | - | - | Yes |
| Self-host | No | No | Yes |
| Open weights | Yes | No | No |
| High warnings | 1 | 1 | 1 |