GA Google Cloud

Gemini Live API on Vertex AI (Gemini Enterprise Agent Platform)

The same Gemini Live models served from Google Cloud with IAM auth, regional endpoints, CMEK, provisioned throughput and live avatar video output. Suits enterprises that need Google Cloud compliance and higher concurrency.

Est. per minute$0.022 - 0.12
2 high-severity warnings

Overview

Best for: Enterprises on Google Cloud that need IAM, CMEK, VPC-SC, provisioned throughput, high concurrency (1,000 sessions per project) or avatar video.

At a glance

Audio in $/1M tok$3
Audio out $/1M tok$12
Free tierNo
Native S2SYes
ToolsYes
Image inYes
Own LLMNo
Voices30
Languages24
Context tokens128,000
Max session min10
Concurrency1,000
WebRTCNo
WebSocketYes
Phone / SIPNo
HIPAAYes
EU dataYes
Open weightsNo

Connections about 10 min; audio-only sessions 15 min without compression. 1,000 concurrent sessions per project on pay-as-you-go. Vertex lists 24 languages. HIPAA BAA for covered Google Cloud services; confirm Live API coverage. Voice count assumed shared with Gemini TTS list. Generic Google Cloud new-customer credits may apply.

Audio in

Raw 16-bit PCM 16 kHz little-endian; JPEG images/video at 1 fps; text.

Audio out

Raw 16-bit PCM 24 kHz little-endian; text; mp4 video for live avatars.

Languages

Vertex overview states 24 languages for multilingual support (Developer API docs claim more); verify per language.

Voices

Same Gemini prebuilt voice set (30 voices on the TTS list); not separately documented for Vertex in the pages read.

Latency

No numeric vendor claim found.

Regions

Regional WebSocket endpoints ({LOCATION}-aiplatform.googleapis.com); CMEK serving profiles in us and eu multi-regions. Check the model's location list before choosing a region.

Compliance

Google Cloud terms; CMEK in us/eu multi-regions; customer data not used for training under Google Cloud terms. HIPAA BAA available for covered Google Cloud services (confirm Live API coverage).

Features

  • native audio, barge-in, VAD
  • affective dialog and proactive audio
  • function calling and Google Search grounding
  • audio transcriptions
  • live avatars (video output) on gemini-3.8-live
  • context window compression and session resumption
  • Provisioned Throughput for Live API
  • CMEK in us and eu multi-regions

Pricing

WhatPriceUnitNotes
Gemini 3.8 Live text input$0.75per 1M tokensnon-global
Gemini 3.8 Live video / image input$1.00per 1M tokens258 tokens per image
Gemini 3.8 Live audio input$3.00per 1M tokens
Gemini 3.8 Live text output (response and reasoning)$4.50per 1M tokens
Gemini 3.8 Live audio output$12.00per 1M tokens
Gemini 3.8 Live avatar video output$1.00per 1M tokens6,192 tokens per second of video while the avatar speaks (about $0.37/min); idle listening not billed
Gemini 2.5 Flash Live API$0.50 text in, $3.00 audio in, $3.00 video/image in, $2.00 text out, $12.00 audio outper 1M tokens
Audio-to-text transcriptiontext output rateper 1M tokenscharged for transcript tokens when transcription is enabled
How the per-minute estimate was worked out

Low = 1 min user audio + 1 min model audio at $3 / $12 per 1M, single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Google states tokens from past turns are re-processed and billed every turn up to the context window limit. Live avatar video adds about $0.37 per minute of avatar speech.

Audio token rate

Audio 25 tokens per second (input and output) = 1,500 tokens/min; image 258 tokens; avatar video 6,192 tokens per second (Vertex Live API billing details).

Free tier: No Live-specific free tier; standard Google Cloud new-customer credits apply.

Source: cloud.google.com

Setup

  1. Create or pick a Google Cloud project with billing enabled and enable the Vertex AI API.
  2. Grant your service account the Vertex AI User role; authenticate with Application Default Credentials.
  3. Install google-genai and create the client with vertexai=True, project and a supported location.
  4. Connect with client.aio.live.connect(model='gemini-3.8-live', config=...) and enable context window compression and session resumption.
  5. For browsers, proxy audio through your backend (or an ADK / partner media server); do not ship service account credentials.
  6. For higher, guaranteed capacity buy Provisioned Throughput for Live API.

Endpoint

wss://{LOCATION}-aiplatform.googleapis.com/ws/google.cloud.aiplatform.v1.LlmBidiService/BidiGenerateContent

Authentication

Google Cloud IAM OAuth access token (ADC / service account) on your backend. No browser-safe key; keep the WebSocket on the server or behind a relay.

Quick start python

import asyncio
from google import genai
from google.genai import types

# Auth: Application Default Credentials (gcloud auth application-default login or a service account)
client = genai.Client(vertexai=True, project="YOUR_PROJECT", location="us-central1")  # pick a region that lists the model
config = types.LiveConnectConfig(
    response_modalities=["AUDIO"],
    system_instruction="You are a friendly support agent. Keep answers short.",
    context_window_compression=types.ContextWindowCompressionConfig(sliding_window=types.SlidingWindow()),
    session_resumption=types.SessionResumptionConfig(),
)

async def run(mic_chunks, play, stop_playback):
    async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
        async def pump():
            async for chunk in mic_chunks():  # raw 16 kHz mono PCM16 bytes
                await session.send_realtime_input(
                    audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000"))
        asyncio.create_task(pump())
        async for msg in session.receive():
            if msg.data:
                play(msg.data)  # 24 kHz mono PCM16
            if msg.server_content and msg.server_content.interrupted:
                stop_playback()  # user barged in
            if msg.go_away:
                print("connection ending in", msg.go_away.time_left)  # reconnect with the resumption handle

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Context re-billing is explicit

Vertex documents that every turn bills all tokens in the session context window, including all previous turns. Long sessions with resumption and compression still pay for whatever stays in the window.

Session clocks

~10 minute connections and 15 minute audio-only sessions (2 minutes with video) unless compression is on. Implement GoAway handling and resumption before going live.

GA vs preview labels disagree

Model docs mark gemini-3.8-live GA while the pricing page still footnotes 'Live API is in Preview'. Confirm SLA coverage with Google before committing to uptime targets.

Avatar video is expensive

Avatar output is 6,192 tokens per second; at $1/1M that is about $0.37 per minute of speaking time, roughly 15x the audio cost of a turn.

Docs moved

Vertex generative AI docs now redirect into the Gemini Enterprise Agent Platform site and the pricing page is titled Agent Platform Pricing. Old bookmarks and blog links may 404.

Language support is narrower than the Developer API docs suggest

Vertex lists 24 languages for Live; test before promising other languages.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Up to 1,000 concurrent sessions per project on pay-as-you-go (not applied to Provisioned Throughput)
  • Connection lifetime around 10 minutes
  • Audio-only sessions 15 minutes and audio-video 2 minutes without context window compression
  • Context window limit 128k tokens; compression trigger 5,000 to 128,000 tokens
  • Session resumption within 24 hours (state typically kept around 10 minutes after disconnect per the same page)
  • Global region is not supported for Live API CMEK (us and eu multi-regions only)

Models and products

NameStatusNotes
gemini-3.8-liveGARecommended. Native audio, transcriptions, VAD, affective dialog, proactive audio, tool use and live avatars.
gemini-live-2.5-flash-native-audioGAPrevious generation; same features minus live avatars. Migration guide to 3.8 Live exists.
gemini-3.5-transcribe-live-previewPreviewSpeech-to-text only with interim/final transcripts and timestamps.

Docs and sources

Docs

Sources used

Not fully verified

Exact list of regions serving gemini-3.8-live (example code uses us-central1 as a placeholder); whether Vertex Live supports ephemeral/browser tokens; voice list on Vertex; per-region price differences (the 3.8 Live rows are labelled Non-global).

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?