Gemini Live API on Vertex AI (Gemini Enterprise Agent Platform)
The same Gemini Live models served from Google Cloud with IAM auth, regional endpoints, CMEK, provisioned throughput and live avatar video output. Suits enterprises that need Google Cloud compliance and higher concurrency.
Overview
Best for: Enterprises on Google Cloud that need IAM, CMEK, VPC-SC, provisioned throughput, high concurrency (1,000 sessions per project) or avatar video.
At a glance
Connections about 10 min; audio-only sessions 15 min without compression. 1,000 concurrent sessions per project on pay-as-you-go. Vertex lists 24 languages. HIPAA BAA for covered Google Cloud services; confirm Live API coverage. Voice count assumed shared with Gemini TTS list. Generic Google Cloud new-customer credits may apply.
Raw 16-bit PCM 16 kHz little-endian; JPEG images/video at 1 fps; text.
Raw 16-bit PCM 24 kHz little-endian; text; mp4 video for live avatars.
Vertex overview states 24 languages for multilingual support (Developer API docs claim more); verify per language.
Same Gemini prebuilt voice set (30 voices on the TTS list); not separately documented for Vertex in the pages read.
No numeric vendor claim found.
Regional WebSocket endpoints ({LOCATION}-aiplatform.googleapis.com); CMEK serving profiles in us and eu multi-regions. Check the model's location list before choosing a region.
Google Cloud terms; CMEK in us/eu multi-regions; customer data not used for training under Google Cloud terms. HIPAA BAA available for covered Google Cloud services (confirm Live API coverage).
Features
- native audio, barge-in, VAD
- affective dialog and proactive audio
- function calling and Google Search grounding
- audio transcriptions
- live avatars (video output) on gemini-3.8-live
- context window compression and session resumption
- Provisioned Throughput for Live API
- CMEK in us and eu multi-regions
Pricing
| What | Price | Unit |
|---|---|---|
| Gemini 3.8 Live text input | $0.75 | per 1M tokens |
| Gemini 3.8 Live video / image input | $1.00 | per 1M tokens |
| Gemini 3.8 Live audio input | $3.00 | per 1M tokens |
| Gemini 3.8 Live text output (response and reasoning) | $4.50 | per 1M tokens |
| Gemini 3.8 Live audio output | $12.00 | per 1M tokens |
| Gemini 3.8 Live avatar video output | $1.00 | per 1M tokens |
| Gemini 2.5 Flash Live API | $0.50 text in, $3.00 audio in, $3.00 video/image in, $2.00 text out, $12.00 audio out | per 1M tokens |
| Audio-to-text transcription | text output rate | per 1M tokens |
Low = 1 min user audio + 1 min model audio at $3 / $12 per 1M, single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Google states tokens from past turns are re-processed and billed every turn up to the context window limit. Live avatar video adds about $0.37 per minute of avatar speech.
Audio 25 tokens per second (input and output) = 1,500 tokens/min; image 258 tokens; avatar video 6,192 tokens per second (Vertex Live API billing details).
Free tier: No Live-specific free tier; standard Google Cloud new-customer credits apply.
Source: cloud.google.com
Setup
- Create or pick a Google Cloud project with billing enabled and enable the Vertex AI API.
- Grant your service account the Vertex AI User role; authenticate with Application Default Credentials.
- Install google-genai and create the client with vertexai=True, project and a supported location.
- Connect with client.aio.live.connect(model='gemini-3.8-live', config=...) and enable context window compression and session resumption.
- For browsers, proxy audio through your backend (or an ADK / partner media server); do not ship service account credentials.
- For higher, guaranteed capacity buy Provisioned Throughput for Live API.
Endpoint
wss://{LOCATION}-aiplatform.googleapis.com/ws/google.cloud.aiplatform.v1.LlmBidiService/BidiGenerateContent
Authentication
Google Cloud IAM OAuth access token (ADC / service account) on your backend. No browser-safe key; keep the WebSocket on the server or behind a relay.
Quick start python
import asyncio
from google import genai
from google.genai import types
# Auth: Application Default Credentials (gcloud auth application-default login or a service account)
client = genai.Client(vertexai=True, project="YOUR_PROJECT", location="us-central1") # pick a region that lists the model
config = types.LiveConnectConfig(
response_modalities=["AUDIO"],
system_instruction="You are a friendly support agent. Keep answers short.",
context_window_compression=types.ContextWindowCompressionConfig(sliding_window=types.SlidingWindow()),
session_resumption=types.SessionResumptionConfig(),
)
async def run(mic_chunks, play, stop_playback):
async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
async def pump():
async for chunk in mic_chunks(): # raw 16 kHz mono PCM16 bytes
await session.send_realtime_input(
audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000"))
asyncio.create_task(pump())
async for msg in session.receive():
if msg.data:
play(msg.data) # 24 kHz mono PCM16
if msg.server_content and msg.server_content.interrupted:
stop_playback() # user barged in
if msg.go_away:
print("connection ending in", msg.go_away.time_left) # reconnect with the resumption handle
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Context re-billing is explicit
Vertex documents that every turn bills all tokens in the session context window, including all previous turns. Long sessions with resumption and compression still pay for whatever stays in the window.
Session clocks
~10 minute connections and 15 minute audio-only sessions (2 minutes with video) unless compression is on. Implement GoAway handling and resumption before going live.
GA vs preview labels disagree
Model docs mark gemini-3.8-live GA while the pricing page still footnotes 'Live API is in Preview'. Confirm SLA coverage with Google before committing to uptime targets.
Avatar video is expensive
Avatar output is 6,192 tokens per second; at $1/1M that is about $0.37 per minute of speaking time, roughly 15x the audio cost of a turn.
Docs moved
Vertex generative AI docs now redirect into the Gemini Enterprise Agent Platform site and the pricing page is titled Agent Platform Pricing. Old bookmarks and blog links may 404.
Language support is narrower than the Developer API docs suggest
Vertex lists 24 languages for Live; test before promising other languages.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Up to 1,000 concurrent sessions per project on pay-as-you-go (not applied to Provisioned Throughput)
- Connection lifetime around 10 minutes
- Audio-only sessions 15 minutes and audio-video 2 minutes without context window compression
- Context window limit 128k tokens; compression trigger 5,000 to 128,000 tokens
- Session resumption within 24 hours (state typically kept around 10 minutes after disconnect per the same page)
- Global region is not supported for Live API CMEK (us and eu multi-regions only)
Models and products
| Name | Status |
|---|---|
| gemini-3.8-live | GA |
| gemini-live-2.5-flash-native-audio | GA |
| gemini-3.5-transcribe-live-preview | Preview |
Docs and sources
Docs
Sources used
- docs.cloud.google.com/vertex-ai/generative-ai/docs/live-api
- docs.cloud.google.com/vertex-ai/generative-ai/docs/live-api/start-manage-session
- cloud.google.com/vertex-ai/generative-ai/pricing
Exact list of regions serving gemini-3.8-live (example code uses us-central1 as a placeholder); whether Vertex Live supports ephemeral/browser tokens; voice list on Vertex; per-region price differences (the 3.8 Live rows are labelled Non-global).