Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Unreal Speech
Unreal Speech
Inworld TTS (Realtime TTS-2, TTS-2 Flash)
Inworld AI
Google Cloud Text-to-Speech (Chirp 3 HD and Gemini-TTS)
Google Cloud
CategoryText-to-speechText-to-speechText-to-speech
StatusGAGAGA
Est. per minute$0.0072 - 0.015$0.0063 - 0.022$0.009 - 0.03
How that was worked out900 chars/min at plan-included rates: Enterprise $8/1M to Basic $16.3/1M (list price). Vendor itself assumes 750 chars/min.900 chars/min. Low = TTS-2 Flash on Growth ($7/1M); high = TTS-2 on On-Demand ($25/1M).Chirp 3 HD: 900 chars x $30/1M = $0.027. Gemini: 60 s x 25 tokens = 1,500 audio tokens/min; 3.8 Flash-Lite $0.009, 3.8 Flash $0.0135 (promo), 2.5 Flash $0.015, 2.5 Pro $0.03 (text input cost negligible).
Pricing modelsubscriptionper-characterper-character
Free tier250K characters/monthOn-Demand: up to 70 minutes of TTSChirp 3 HD: 1M characters/month; WaveNet/Standard 4M; Gemini-TTS: none
Connects byHTTP chunkedWebSocket, HTTP chunkedgRPC, HTTP chunked
Audio inTextText with inline tags (pauses, pronunciation, steering on TTS-2)Text, SSML (legacy voices), natural-language prompt for Gemini-TTS
Audio outMP3 (bitrate 16k-320k, default 192k); other formats not verifiedLINEAR16 (WAV header per chunk), PCM, MP3 (default), OGG_OPUS, ALAW, MULAW, WAV; 8-48 kHz (default 48 kHz)Streaming: PCM (default), ALAW, MULAW, OGG_OPUS. Batch: LINEAR16, ALAW, MULAW, MP3, OGG_OPUS, PCM.
LanguagesEnglish plus Chinese, Spanish, French, Hindi, Italian voices (V8 voice list)Docs say 200+ languages and locales; release notes describe 15 production-quality plus 90+ experimental for TTS-2Chirp 3 HD: 53 locales; Gemini-TTS: see per-model list
Latency (vendor claim)Vendor: /stream returns audio in ~0.3 s.Vendor: P90 server-side TTFB 100 ms (TTS-2), 20 ms (TTS-2 Flash).No numeric claim on the pages checked; Gemini-TTS described as 'very low latency'.
Key limits
  • /stream 1,000 chars per request
  • /speech 3,000 chars
  • /synthesisTasks 500,000 chars
  • Concurrent requests: On-Demand 5, Creator 10, Builder 50, Developer 150, Growth 500
  • 2,000 UTF-16 code units per send_text message
  • Socket closes after 10 min of inactivity across all contexts
  • Sync 2,000 chars, HTTP streaming 4,000 chars, async 100,000 chars per request
  • Gemini-TTS: 8,192 input tokens, 16,384 output tokens per request
  • StreamingSynthesize: first message must be config only; Preview (Pre-GA terms)
  • Quotas per project; see quotas page
High-severity warnings
  • None
  • None
  • Bidi streaming only for Chirp 3 HD and still Preview
  • Gemini 3.8 TTS is not on the Cloud TTS API
ComplianceNot verified.Zero data retention supported (models page). Certifications not re-verified.Google Cloud data terms and regional endpoints; certifications not re-verified here.
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/1M chars$16.33$25$30
Free tierYesYesYes
Free credit $---
Free tier commercial---
Voices--30
Cloning-YesYes
Instant clone--Yes
Text stream inNoYesYes
TimestampsYesYes-
Emotion-YesYes
SSML--Yes
8 kHz phone-YesYes
Latency ms300100-
Languages615-
Max session min---
Concurrency-5-
WebRTCNoNoNo
WebSocketNoYesNo
gRPCNoNoYes
HIPAA---
SOC 2---
EU data--Yes
Self-hostNoNoNo
Open weightsNoNoNo
High warnings002