Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Gemini Live API (Gemini Developer API / Google AI Studio)
Google
Gemini Live API on Vertex AI (Gemini Enterprise Agent Platform)
Google Cloud
Amazon Nova 2 Sonic / Nova 2.5 Sonic (Bedrock)
Amazon Web Services
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minute$0.022 - 0.12$0.022 - 0.12$0.022 - 0.12
How that was worked outLow = gemini-3.8-live, 1 min user audio (1,500 tokens x $3/1M) + 1 min model audio (1,500 tokens x $12/1M), single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Live models do not support context caching, so there is no cached discount. Video input or thinking tokens add more.Low = 1 min user audio + 1 min model audio at $3 / $12 per 1M, single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Google states tokens from past turns are re-processed and billed every turn up to the context window limit. Live avatar video adds about $0.37 per minute of avatar speech.UNVERIFIED token rate. Assuming ~25 speech tokens/s (community figure): low = 1 min user speech (1,500 x $3/1M) + 1 min model speech (1,500 x $12/1M) = $0.0225, single turn. High uses the same 10 minute, 50 turn model as other providers and assumes history is re-processed each turn like other S2S APIs (AWS does not document this). Third-party sites quote roughly $0.015/min.
Pricing modelper-tokenper-tokenper-token
Free tierYes: free tier for all Live models with lower rate limits; free-tier content may be used to improve Google products.No Live-specific free tier; standard Google Cloud new-customer credits apply.None specific to Nova Sonic (AWS Free Tier credits may apply to new accounts).
Connects byWebSocketWebSocketHTTP/2 bidirectional (InvokeModelWithBidirectionalStream)
Audio inRaw 16-bit PCM little-endian, natively 16 kHz (other rates resampled if the MIME type says so, e.g. audio/pcm;rate=16000). Images/video as JPEG or PNG frames, max 1 fps.Raw 16-bit PCM 16 kHz little-endian; JPEG images/video at 1 fps; text.audio/lpcm 16-bit mono at 8, 16 or 24 kHz, base64 in audioInput events (~32 ms frames streamed in real time).
Audio outRaw 16-bit PCM little-endian at 24 kHz (always).Raw 16-bit PCM 24 kHz little-endian; text; mp4 video for live avatars.audio/lpcm 16-bit mono at 8, 16 or 24 kHz, base64 audioOutput events.
LanguagesLive guide lists 99 languages (overview page says 70); native audio models pick the language automatically and do not accept a language code.Vertex overview states 24 languages for multilingual support (Developer API docs claim more); verify per language.English (US, UK, India, Australia), French, Italian, German, Spanish, Portuguese, Hindi, with automatic language detection and switching.
Latency (vendor claim)No numeric vendor claim found.No numeric vendor claim found.Vendor claim: March 2026 refresh cut user-perceived p50 latency by 150 ms; Nova 2.5 Sonic claimed lower latency than Nova 2 Sonic (no absolute numbers published).
Key limits
  • Connection lifetime around 10 minutes (GoAway with timeLeft is sent before close)
  • Without compression: audio-only sessions 15 minutes, audio+video 2 minutes
  • Context window 128k tokens for native audio models (3.8 Live lists 131,072 input)
  • Session resumption tokens valid 2 hours after the last session ends (Developer API)
  • Up to 1,000 concurrent sessions per project on pay-as-you-go (not applied to Provisioned Throughput)
  • Connection lifetime around 10 minutes
  • Audio-only sessions 15 minutes and audio-video 2 minutes without context window compression
  • Context window limit 128k tokens; compression trigger 5,000 to 128,000 tokens
  • Connection limit 8 minutes; renew the connection and carry history forward (AWS samples and Strands Bidi Agents do this)
  • Default quota 20 concurrent InvokeModelWithBidirectionalStream sessions per account per Region for Nova 2 Sonic, documented as not adjustable
  • Context window: Nova 2 Sonic 1M tokens (64K output); Nova 2.5 Sonic 256K per launch post
  • Conversation history can only be injected once, after the system prompt and before audio streaming
High-severity warnings
  • 10 minute connections, 15 minute sessions
  • Whole context re-billed every turn, no caching
  • Free tier trains on your data
  • Context re-billing is explicit
  • Session clocks
  • 20 concurrent sessions, not adjustable
  • 8 minute connection limit
CompliancePaid tier data is not used to improve products; free tier is. No BAA or data residency on the Developer API; use Vertex AI for enterprise compliance (CMEK, VPC-SC, regional processing).Google Cloud terms; CMEK in us/eu multi-regions; customer data not used for training under Google Cloud terms. HIPAA BAA available for covered Google Cloud services (confirm Live API coverage).Bedrock data is not used to train models and stays in the chosen Region (in-Region inference only). Bedrock is HIPAA eligible and covered by AWS SOC reports; confirm Nova Sonic is listed in your BAA scope.
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min---
Audio in $/1M tok$3$3$3
Audio out $/1M tok$12$12$12
Free tierYesNoNo
Free credit $---
Native S2SYesYesYes
ToolsYesYesYes
Image inYesYes-
Own LLMNoNoNo
Voices303016
Cloning---
Latency ms---
Languages99247
Context tokens131,072128,0001,000,000
Max session min10108
Concurrency-1,00020
WebRTCNoNoNo
WebSocketYesYesNo
Phone / SIPNoNoNo
HIPAANoYesYes
SOC 2--Yes
EU dataNoYesYes
Open weightsNoNoNo
High warnings322