Gemini Live API (Gemini Developer API / Google AI Studio)
Stateful WebSocket API for Gemini native-audio models with audio, text and 1 fps image/video input and spoken output. Much cheaper per token than OpenAI, with a free tier; suits prototypes and cost-sensitive voice agents that can live with short connection limits.
Overview
Best for: Low-cost voice agents, prototypes on the free tier, and multimodal assistants that also look at a camera or screen (1 fps).
At a glance
Connections drop at about 10 min; audio-only sessions stop at 15 min (2 min with video) without compression and resumption. 99 languages per the Live guide (overview page says 70). Free tier data is used to improve Google products. No BAA or region choice on the Developer API.
Raw 16-bit PCM little-endian, natively 16 kHz (other rates resampled if the MIME type says so, e.g. audio/pcm;rate=16000). Images/video as JPEG or PNG frames, max 1 fps.
Raw 16-bit PCM little-endian at 24 kHz (always).
Live guide lists 99 languages (overview page says 70); native audio models pick the language automatically and do not accept a language code.
30 prebuilt voices shared with Gemini TTS (e.g. Kore, Puck, Zephyr).
No numeric vendor claim found.
No region selection on the Developer API (Google-managed global serving). Use Vertex AI for regional control and CMEK.
Paid tier data is not used to improve products; free tier is. No BAA or data residency on the Developer API; use Vertex AI for enterprise compliance (CMEK, VPC-SC, regional processing).
Features
- native audio output
- barge-in (serverContent.interrupted)
- automatic VAD with start/end sensitivity, prefix padding, silence duration; or manual activityStart/activityEnd
- function calling (async/NON_BLOCKING on 3.8) and Google Search grounding
- input and output audio transcription
- affective dialog (v1beta, not on 3.1 Flash Live)
- proactive audio (model may choose not to answer)
- thinking on extended-thinking and 3.1 Flash Live models
- image and video frame input (1 fps)
- context window compression (sliding window)
- session resumption handles and GoAway warnings
- ephemeral tokens with config locking (liveConnectConstraints)
Pricing
| What | Price | Unit |
|---|---|---|
| gemini-3.8-live (and extended thinking, 3.1 Flash Live) text input | $0.75 | per 1M tokens |
| audio input | $3.00 | per 1M tokens |
| image / video input | $1.00 | per 1M tokens |
| text output (incl. thinking) | $4.50 | per 1M tokens |
| audio output | $12.00 | per 1M tokens |
| gemini-2.5-flash-native-audio-preview-12-2025 | $0.50 text in, $3.00 audio/video in, $2.00 text out, $12.00 audio out | per 1M tokens |
| gemini-3.5-live-translate-preview | $3.50 in / $21.00 out | per 1M audio tokens |
| gemini-3.5-transcribe-live | $3.50 audio in / $21.00 text out | per 1M tokens |
| Google Search grounding (Gemini 3+) | 5,000 free requests/month, then $14 | per 1,000 search queries |
| Free tier | $0 |
Low = gemini-3.8-live, 1 min user audio (1,500 tokens x $3/1M) + 1 min model audio (1,500 tokens x $12/1M), single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Live models do not support context caching, so there is no cached discount. Video input or thinking tokens add more.
25 tokens per second of audio in and out = 1,500 tokens/min (Google pricing footnotes and Vertex Live billing details); 258 tokens per image.
Free tier: Yes: free tier for all Live models with lower rate limits; free-tier content may be used to improve Google products.
Source: ai.google.dev
Setup
- Sign in to Google AI Studio (aistudio.google.com) and create an API key for a Google Cloud project.
- Start on the free tier; enable billing on the project to move to paid tiers and keep data out of product improvement.
- Install the SDK (npm i @google/genai or pip install google-genai).
- Server-to-server: connect with the API key. Client-to-server: your server calls client.authTokens.create (v1beta) with uses: 1 and liveConnectConstraints, and the browser uses token.name as its apiKey.
- Enable contextWindowCompression and sessionResumption from day one and handle goAway by reconnecting with the latest handle.
- Stream 16 kHz PCM, play 24 kHz PCM, and stop playback on serverContent.interrupted.
Endpoint
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent?key=API_KEY (ephemeral tokens use BidiGenerateContentConstrained with access_token)
Authentication
API key on the server only. Browsers and mobile use ephemeral tokens from POST https://generativelanguage.googleapis.com/v1beta/auth_tokens, sent as access_token query parameter or Authorization: Token <token>; lock model and config with liveConnectConstraints.
Quick start javascript
import { GoogleGenAI, Modality } from "@google/genai";
// Server side with an API key. In a browser, create an ephemeral token on your server
// (client.authTokens.create) and pass token.name as apiKey instead.
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const session = await ai.live.connect({
model: "gemini-3.8-live",
config: {
responseModalities: [Modality.AUDIO],
systemInstruction: "You are a friendly support agent. Keep answers short.",
speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Kore" } } },
contextWindowCompression: { slidingWindow: {} }, // lifts the 15 min audio session cap
sessionResumption: {}, // survive the ~10 min connection limit
outputAudioTranscription: {},
},
callbacks: {
onmessage: (msg) => {
for (const part of msg.serverContent?.modelTurn?.parts ?? []) // a message can hold several parts
if (part.inlineData?.data) playPcm24k(Buffer.from(part.inlineData.data, "base64"));
if (msg.serverContent?.interrupted) stopPlayback(); // barge-in
if (msg.sessionResumptionUpdate?.newHandle) saveHandle(msg.sessionResumptionUpdate.newHandle);
if (msg.goAway) scheduleReconnect(msg.goAway.timeLeft);
},
onerror: (e) => console.error(e),
onclose: (e) => console.log("closed", e.reason),
},
});
// 16 kHz mono PCM16 chunks from the microphone
export const sendAudio = (pcm16) =>
session.sendRealtimeInput({ audio: { data: pcm16.toString("base64"), mimeType: "audio/pcm;rate=16000" } });
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
10 minute connections, 15 minute sessions
Connections drop around 10 minutes and audio-only sessions stop at 15 minutes (2 minutes with video) unless you enable context window compression and session resumption. Every production app needs reconnect logic driven by goAway.
Whole context re-billed every turn, no caching
Google bills all tokens in the session context window on each turn and Live models do not support caching. At 25 tokens/s, long calls grow fast; use a sliding window with a sensible trigger_tokens.
Free tier trains on your data
Free-tier content is used to improve Google products. Do not run real customer calls on a free-tier key.
Model churn
Google moved from 2.0 Flash Live and Live 2.5 Flash (shut down Dec 9, 2025) to 2.5 native audio, 3.1 Flash Live preview (Mar 2026) and 3.8 Live (Sept 2026). The 2.5 models are now restricted to previous users. Config differences (thinkingLevel rejected on 3.8, async tools default) break copy-pasted setups.
Audio-only responses
Native audio models only support the AUDIO response modality. For text you must enable output audio transcription, and a single server event can carry several parts, so loop over all parts.
Manual VAD is unforgiving
With automatic VAD off the server adds no pre-speech buffer or silence tolerance; keep at least 500 ms end-of-speech silence and send audioStreamEnd when the mic pauses for over a second.
Ordering vs responsiveness
send_realtime_input favours speed over strict ordering; use send_client_content when order matters. On 3.8 models turn_complete=true in send_client_content interrupts generation.
No echo cancellation server side
Use browser echoCancellation or headphones; otherwise the model hears itself and barges in on its own speech.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Connection lifetime around 10 minutes (GoAway with timeLeft is sent before close)
- Without compression: audio-only sessions 15 minutes, audio+video 2 minutes
- Context window 128k tokens for native audio models (3.8 Live lists 131,072 input)
- Session resumption tokens valid 2 hours after the last session ends (Developer API)
- Ephemeral tokens: new-session window default 1 minute, message window default 30 minutes
- Rate limits are per project and shown in AI Studio; Live concurrency is not published on the rate-limit page
Models and products
| Name | Status |
|---|---|
| gemini-3.8-live | GA |
| gemini-3.8-live-extended-thinking | GA |
| gemini-3.1-flash-live-preview | Preview |
| gemini-2.5-flash-native-audio-preview-12-2025 | Preview |
| gemini-3.5-live-translate-preview | Preview |
| gemini-3.5-transcribe-live | Preview |
| gemini-2.0-flash-live-001, gemini-live-2.5-flash-preview | Deprecated |
Docs and sources
Docs
- Live API overview
- Live capabilities guide
- Session management
- Ephemeral tokens
- WebSocket get started
- Pricing
Sources used
- ai.google.dev/gemini-api/docs/pricing
- ai.google.dev/gemini-api/docs/models
- ai.google.dev/gemini-api/docs/models/gemini-3.8-live
- ai.google.dev/gemini-api/docs/live
- ai.google.dev/gemini-api/docs/live-guide
- ai.google.dev/gemini-api/docs/live-session
- ai.google.dev/gemini-api/docs/ephemeral-tokens
- ai.google.dev/gemini-api/docs/live-api/get-started-websocket
- ai.google.dev/gemini-api/docs/changelog
- ai.google.dev/gemini-api/docs/speech-generation
- cloud.google.com/vertex-ai/generative-ai/pricing
Live API concurrent session limits per tier on the Developer API; exact language count (99 vs 70 on different pages); whether 'half-cascade' models still exist (current docs only describe native audio); latency figures. The live guide still says 'the Live API is in preview' while the models are labelled stable/GA.