Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Amazon Polly (generative and bidirectional streaming) Amazon Web Services | Mistral Voxtral TTS Mistral AI | OpenAI Text-to-Speech (gpt-4o-mini-tts) OpenAI | |
|---|---|---|---|
| Category | Text-to-speech | Text-to-speech | Text-to-speech |
| Status | GA | GA | GA |
| Est. per minute | $0.014 - 0.027 | $0.014 | $0.013 - 0.027 |
| How that was worked out | 900 chars/min. Bidi streaming requires generative ($30/1M = $0.027); neural $16/1M = $0.0144 without input streaming. | 900 chars/min x $16/1M | 900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now). |
| Pricing model | per-character | per-character | per-token |
| Free tier | 12 months: 5M standard, 1M neural, 500K long-form, 100K generative characters per month; new accounts also get up to $200 AWS credits | Not verified | None specific to TTS |
| Connects by | HTTP chunked, HTTP/2 bidirectional event stream | HTTP chunked | HTTP chunked, SSE |
| Audio in | Plain text or SSML | Text | Text plus optional free-text instructions |
| Audio out | mp3, ogg_opus, ogg_vorbis, pcm (stream API; JSON speech marks not supported on the stream API) | Not verified | mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless) |
| Languages | Stream API LanguageCode list covers ~40 locales (only needed for bilingual voices) | 9: en, fr, es, pt, it, nl, de, ar, hi | Follows Whisper language support; voices optimised for English |
| Latency (vendor claim) | No numeric claim found. | Model card (self-hosted, 1 GPU): 70 ms at concurrency 1, 331 ms at 16, 552 ms at 32. | No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | AWS compliance programs (GovCloud availability for standard/neural). | Not verified. | OpenAI platform terms; disclosure of AI voice required by usage policy. |
| Self-hostable | No | Yes | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| $/1M chars | $30 | $16 | $15 |
| Free tier | Yes | - | No |
| Free credit $ | $200 | - | - |
| Free tier commercial | - | - | - |
| Voices | - | - | - |
| Cloning | No | Yes | Yes |
| Instant clone | No | - | - |
| Text stream in | Yes | - | No |
| Timestamps | Yes | - | - |
| Emotion | - | - | Yes |
| SSML | Yes | - | - |
| 8 kHz phone | No | - | No |
| Latency ms | - | 70 | - |
| Languages | - | 9 | - |
| Max session min | - | - | - |
| Concurrency | - | - | - |
| WebRTC | No | No | No |
| WebSocket | No | No | No |
| gRPC | No | No | No |
| HIPAA | - | - | - |
| SOC 2 | - | - | - |
| EU data | Yes | - | - |
| Self-host | No | Yes | No |
| Open weights | No | Yes | No |
| High warnings | 1 | 1 | 1 |