Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Soniox Text-to-Speech
Soniox
ElevenLabs TTS API
ElevenLabs
OpenAI Text-to-Speech (gpt-4o-mini-tts)
OpenAI
CategoryText-to-speechText-to-speechText-to-speech
StatusGAGAGA
Est. per minute$0.012$0.0099 - 0.072$0.013 - 0.027
How that was worked outVendor ~$0.70/hour = $0.0117/min900 chars/min. Low = v4 Turbo promo rate ($0.011/1K, ends Oct 12 2026; $0.036/min after). Flash v2.5 = $0.036/min. High = v3 or Multilingual v2 at $0.08/1K.900 chars/min for tts-1 ($15/1M) and tts-1-hd ($30/1M). gpt-4o-mini-tts is token-billed; OpenAI no longer shows a per-minute estimate on the pricing page (an earlier page estimated about $0.015/min, unverified now).
Pricing modelper-tokenper-characterper-token
Free tierNot verifiedFree / pay-as-you-go: 10,000-20,000 characters depending on model. Commercial-use terms on the free tier not re-verified; check the plan terms.None specific to TTS
Connects byWebSocket, HTTP chunkedWebSocket, HTTP chunkedHTTP chunked, SSE
Audio inText, streamedText (SSML parsing optional via enable_ssml_parsing on the WebSocket)Text plus optional free-text instructions
Audio outNot verifiedMP3 by default; output_format values follow codec_samplerate_bitrate, e.g. mp3_44100_128, pcm_16000/22050/24000/44100, ulaw_8000, alaw_8000, opus_48000_* (format list from third-party mirrors of the API reference; some higher-quality formats are tier-gated)mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE, headerless)
Languages60+ (vendor launch post)Flash v2.5: 32; v3: 70+; v4: 90+Follows Whisper language support; voices optimised for English
Latency (vendor claim)No numeric claim verified.Vendor claims: Flash v2.5 ~75 ms model latency, v4 Turbo ~100 ms median inference, v3 conversational ~280 ms. All exclude network and application latency.No numeric TTFB claim on the guide; WAV/PCM recommended for fastest first bytes.
Key limits
  • Not published for TTS
  • Chars per request: Flash v2.5 40,000; Multilingual v2 and v4 10,000; v3 5,000 (models page)
  • WebSocket inactivity_timeout default 20 s, max 180 s
  • Third-party reports (Vapi support): max 5 simultaneous contexts per multi-context WebSocket; not confirmed on official pages
  • Plan concurrency limits are not shown on the API pricing page; third-party lists (Free 2 ... Business 15) are unofficial
  • gpt-4o-mini-tts max 2,000 input tokens per request
  • Rate limits by tier: Build 2,000 RPM / 150K TPM; Launch 10,000 RPM / 2M TPM; Grow 10,000 RPM / 8M TPM
High-severity warnings
  • None
  • v3 and v4 do not work on the classic TTS WebSocket
  • v4 launch prices are promotional
  • No text-input streaming
ComplianceVendor says generated audio is not stored and not used for training.Data-residency endpoints for EU, India and Singapore exist. Certifications not re-verified in this pass.OpenAI platform terms; disclosure of AI voice required by usage policy.
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/1M chars$13$40$15
Free tier-YesNo
Free credit $---
Free tier commercial---
Voices---
Cloning-YesYes
Instant clone---
Text stream inYesYesNo
TimestampsYesYes-
Emotion--Yes
SSML-Yes-
8 kHz phone-YesNo
Latency ms-75-
Languages6032-
Max session min---
Concurrency-6-
WebRTCNoNoNo
WebSocketYesYesNo
gRPCNoNoNo
HIPAA---
SOC 2---
EU data-Yes-
Self-hostNoNoNo
Open weightsNoNoNo
High warnings021