Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Mistral Voxtral TTS Mistral AI | Amazon Polly (generative and bidirectional streaming) Amazon Web Services | OpenAI Text-to-Speech (gpt-4o-mini-tts) OpenAI | |
|---|---|---|---|
| Category | Text-to-speech | Text-to-speech | Text-to-speech |
| Status | GA | GA | GA |
| Est. per minute | $0.014 | $0.014 - 0.027 | $0.013 - 0.027 |
| How that was worked out | 900 chars/min x $16/1M | 900 chars/min. Bidi streaming requires generative ($30/1M = $0.027); neural $16/1M = $0.0144 without input streaming. | 900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now). |
| Pricing model | per-character | per-character | per-token |
| Free tier | Not verified | 12 months: 5M standard, 1M neural, 500K long-form, 100K generative characters per month; new accounts also get up to $200 AWS credits | None specific to TTS |
| Connects by | HTTP chunked | HTTP chunked, HTTP/2 bidirectional event stream | HTTP chunked, SSE |
| Audio in | Text | Plain text or SSML | Text plus optional free-text instructions |
| Audio out | Not verified | mp3, ogg_opus, ogg_vorbis, pcm (stream API; JSON speech marks not supported on the stream API) | mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless) |
| Languages | 9: en, fr, es, pt, it, nl, de, ar, hi | Stream API LanguageCode list covers ~40 locales (only needed for bilingual voices) | Follows Whisper language support; voices optimised for English |
| Latency (vendor claim) | Model card (self-hosted, 1 GPU): 70 ms at concurrency 1, 331 ms at 16, 552 ms at 32. | No numeric claim found. | No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Not verified. | AWS compliance programs (GovCloud availability for standard/neural). | OpenAI platform terms; disclosure of AI voice required by usage policy. |
| Self-hostable | Yes | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| $/1M chars | $16 | $30 | $15 |
| Free tier | - | Yes | No |
| Free credit $ | - | $200 | - |
| Free tier commercial | - | - | - |
| Voices | - | - | - |
| Cloning | Yes | No | Yes |
| Instant clone | - | No | - |
| Text stream in | - | Yes | No |
| Timestamps | - | Yes | - |
| Emotion | - | - | Yes |
| SSML | - | Yes | - |
| 8 kHz phone | - | No | No |
| Latency ms | 70 | - | - |
| Languages | 9 | - | - |
| Max session min | - | - | - |
| Concurrency | - | - | - |
| WebRTC | No | No | No |
| WebSocket | No | No | No |
| gRPC | No | No | No |
| HIPAA | - | - | - |
| SOC 2 | - | - | - |
| EU data | - | Yes | - |
| Self-host | Yes | No | No |
| Open weights | Yes | No | No |
| High warnings | 1 | 1 | 1 |