Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Soniox Text-to-Speech Soniox | ElevenLabs TTS API ElevenLabs | OpenAI Text-to-Speech (gpt-4o-mini-tts) OpenAI | |
|---|---|---|---|
| Category | Text-to-speech | Text-to-speech | Text-to-speech |
| Status | GA | GA | GA |
| Est. per minute | $0.012 | $0.0099 - 0.072 | $0.013 - 0.027 |
| How that was worked out | Vendor ~$0.70/hour = $0.0117/min | 900 chars/min. Low = v4 Turbo promo rate ($0.011/1K, ends Oct 12 2026; $0.036/min after). Flash v2.5 = $0.036/min. High = v3 or Multilingual v2 at $0.08/1K. | 900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now). |
| Pricing model | per-token | per-character | per-token |
| Free tier | Not verified | Free / pay-as-you-go: 10,000-20,000 characters depending on model. Commercial-use terms on the free tier not re-verified; check the plan terms. | None specific to TTS |
| Connects by | WebSocket, HTTP chunked | WebSocket, HTTP chunked | HTTP chunked, SSE |
| Audio in | Text, streamed | Text (SSML parsing optional via enable_ssml_parsing on the WebSocket) | Text plus optional free-text instructions |
| Audio out | Not verified | MP3 by default; output_format values follow codec_samplerate_bitrate, e.g. mp3_44100_128, pcm_16000/22050/24000/44100, ulaw_8000, alaw_8000, opus_48000_* (format list from third-party mirrors of the API reference; some higher-quality formats are tier-gated) | mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless) |
| Languages | 60+ (vendor launch post) | Flash v2.5: 32; v3: 70+; v4: 90+ | Follows Whisper language support; voices optimised for English |
| Latency (vendor claim) | No numeric claim verified. | Vendor claims: Flash v2.5 ~75 ms model latency, v4 Turbo ~100 ms median inference, v3 conversational ~280 ms. All exclude network and application latency. | No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Vendor says generated audio is not stored and not used for training. | Data-residency endpoints for EU, India and Singapore exist. Certifications not re-verified in this pass. | OpenAI platform terms; disclosure of AI voice required by usage policy. |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| $/1M chars | $13 | $40 | $15 |
| Free tier | - | Yes | No |
| Free credit $ | - | - | - |
| Free tier commercial | - | - | - |
| Voices | - | - | - |
| Cloning | - | Yes | Yes |
| Instant clone | - | - | - |
| Text stream in | Yes | Yes | No |
| Timestamps | Yes | Yes | - |
| Emotion | - | - | Yes |
| SSML | - | Yes | - |
| 8 kHz phone | - | Yes | No |
| Latency ms | - | 75 | - |
| Languages | 60 | 32 | - |
| Max session min | - | - | - |
| Concurrency | - | 6 | - |
| WebRTC | No | No | No |
| WebSocket | Yes | Yes | No |
| gRPC | No | No | No |
| HIPAA | - | - | - |
| SOC 2 | - | - | - |
| EU data | - | Yes | - |
| Self-host | No | No | No |
| Open weights | No | No | No |
| High warnings | 0 | 2 | 1 |