Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Google Cloud Text-to-Speech (Chirp 3 HD and Gemini-TTS)
Google Cloud
Murf Falcon / Falcon 2
Murf AI
ElevenLabs TTS API
ElevenLabs
CategoryText-to-speechText-to-speechText-to-speech
StatusGAGAGA
Est. per minute$0.009 - 0.03$0.009 - 0.01$0.0099 - 0.072
How that was worked outChirp 3 HD: 900 chars x $30/1M = $0.027. Gemini: 60 s x 25 tokens = 1,500 audio tokens/min; 3.8 Flash-Lite $0.009, 3.8 Flash $0.0135 (promo), 2.5 Flash $0.015, 2.5 Pro $0.03 (text input cost negligible).Vendor headline 1 cent/min; 900 chars x $0.01/1K = $0.009900 chars/min. Low = v4 Turbo promo rate ($0.011/1K, ends Oct 12 2026; $0.036/min after). Flash v2.5 = $0.036/min. High = v3 or Multilingual v2 at $0.08/1K.
Pricing modelper-characterper-characterper-character
Free tierChirp 3 HD: 1M characters/month; WaveNet/Standard 4M; Gemini-TTS: noneThird-party sources conflict ($10 monthly credit vs 100K-character trial)Free / pay-as-you-go: 10,000-20,000 characters depending on model. Commercial-use terms on the free tier not re-verified; check the plan terms.
Connects bygRPC, HTTP chunkedWebSocket, HTTP chunkedWebSocket, HTTP chunked
Audio inText, SSML (legacy voices), natural-language prompt for Gemini-TTSTextText (SSML parsing optional via enable_ssml_parsing on the WebSocket)
Audio outStreaming: PCM (default), ALAW, MULAW, OGG_OPUS. Batch: LINEAR16, ALAW, MULAW, MP3, OGG_OPUS, PCM.WAV or PCM (16-bit LE), example sample_rate 24000, MONOMP3 by default; output_format values follow codec_samplerate_bitrate, e.g. mp3_44100_128, pcm_16000/22050/24000/44100, ulaw_8000, alaw_8000, opus_48000_* (format list from third-party mirrors of the API reference; some higher-quality formats are tier-gated)
LanguagesChirp 3 HD: 53 locales; Gemini-TTS: see per-model list35+ (vendor)Flash v2.5: 32; v3: 70+; v4: 90+
Latency (vendor claim)No numeric claim on the pages checked; Gemini-TTS described as 'very low latency'.Vendor: Falcon 55 ms model latency / 130 ms TTFA; Falcon 2 sub-100 ms TTFA.Vendor claims: Flash v2.5 ~75 ms model latency, v4 Turbo ~100 ms median inference, v3 conversational ~280 ms. All exclude network and application latency.
Key limits
  • Gemini-TTS: 8,192 input tokens, 16,384 output tokens per request
  • StreamingSynthesize: first message must be config only; Preview (Pre-GA terms)
  • Quotas per project; see quotas page
  • Streaming concurrency: 5 on US-East, 2 on all other regions (global router follows regional limits)
  • Up to 10x your concurrency in open WebSocket connections
  • Idle sessions closed after 3 minutes
  • Chars per request: Flash v2.5 40,000; Multilingual v2 and v4 10,000; v3 5,000 (models page)
  • WebSocket inactivity_timeout default 20 s, max 180 s
  • Third-party reports (Vapi support): max 5 simultaneous contexts per multi-context WebSocket; not confirmed on official pages
  • Plan concurrency limits are not shown on the API pricing page; third-party lists (Free 2 ... Business 15) are unofficial
High-severity warnings
  • Bidi streaming only for Chirp 3 HD and still Preview
  • Gemini 3.8 TTS is not on the Cloud TTS API
  • Very low default concurrency outside US-East
  • v3 and v4 do not work on the classic TTS WebSocket
  • v4 launch prices are promotional
ComplianceGoogle Cloud data terms and regional endpoints; certifications not re-verified here.Not verified.Data-residency endpoints for EU, India and Singapore exist. Certifications not re-verified in this pass.
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/1M chars$30$10$40
Free tierYes-Yes
Free credit $---
Free tier commercial---
Voices30--
CloningYes-Yes
Instant cloneYes--
Text stream inYesYesYes
Timestamps-YesYes
EmotionYesYes-
SSMLYes-Yes
8 kHz phoneYes-Yes
Latency ms-10075
Languages-3532
Max session min---
Concurrency-56
WebRTCNoNoNo
WebSocketNoYesYes
gRPCYesNoNo
HIPAA---
SOC 2---
EU dataYesYesYes
Self-hostNoNoNo
Open weightsNoNoNo
High warnings212