GA Google

Gemini Live API (Gemini Developer API / Google AI Studio)

Stateful WebSocket API for Gemini native-audio models with audio, text and 1 fps image/video input and spoken output. Much cheaper per token than OpenAI, with a free tier; suits prototypes and cost-sensitive voice agents that can live with short connection limits.

Est. per minute$0.022 - 0.12
3 high-severity warnings

Overview

Best for: Low-cost voice agents, prototypes on the free tier, and multimodal assistants that also look at a camera or screen (1 fps).

At a glance

Audio in $/1M tok$3
Audio out $/1M tok$12
Free tierYes
Native S2SYes
ToolsYes
Image inYes
Own LLMNo
Voices30
Languages99
Context tokens131,072
Max session min10
WebRTCNo
WebSocketYes
Phone / SIPNo
HIPAANo
EU dataNo
Open weightsNo

Connections drop at about 10 min; audio-only sessions stop at 15 min (2 min with video) without compression and resumption. 99 languages per the Live guide (overview page says 70). Free tier data is used to improve Google products. No BAA or region choice on the Developer API.

Audio in

Raw 16-bit PCM little-endian, natively 16 kHz (other rates resampled if the MIME type says so, e.g. audio/pcm;rate=16000). Images/video as JPEG or PNG frames, max 1 fps.

Audio out

Raw 16-bit PCM little-endian at 24 kHz (always).

Languages

Live guide lists 99 languages (overview page says 70); native audio models pick the language automatically and do not accept a language code.

Voices

30 prebuilt voices shared with Gemini TTS (e.g. Kore, Puck, Zephyr).

Latency

No numeric vendor claim found.

Regions

No region selection on the Developer API (Google-managed global serving). Use Vertex AI for regional control and CMEK.

Compliance

Paid tier data is not used to improve products; free tier is. No BAA or data residency on the Developer API; use Vertex AI for enterprise compliance (CMEK, VPC-SC, regional processing).

Features

  • native audio output
  • barge-in (serverContent.interrupted)
  • automatic VAD with start/end sensitivity, prefix padding, silence duration; or manual activityStart/activityEnd
  • function calling (async/NON_BLOCKING on 3.8) and Google Search grounding
  • input and output audio transcription
  • affective dialog (v1beta, not on 3.1 Flash Live)
  • proactive audio (model may choose not to answer)
  • thinking on extended-thinking and 3.1 Flash Live models
  • image and video frame input (1 fps)
  • context window compression (sliding window)
  • session resumption handles and GoAway warnings
  • ephemeral tokens with config locking (liveConnectConstraints)

Pricing

WhatPriceUnitNotes
gemini-3.8-live (and extended thinking, 3.1 Flash Live) text input$0.75per 1M tokenspaid tier
audio input$3.00per 1M tokensGoogle also quotes $0.005/min
image / video input$1.00per 1M tokensor $0.002/min
text output (incl. thinking)$4.50per 1M tokens
audio output$12.00per 1M tokensor $0.018/min
gemini-2.5-flash-native-audio-preview-12-2025$0.50 text in, $3.00 audio/video in, $2.00 text out, $12.00 audio outper 1M tokenslegacy access only
gemini-3.5-live-translate-preview$3.50 in / $21.00 outper 1M audio tokensabout $0.0053/min in and $0.0315/min out
gemini-3.5-transcribe-live$3.50 audio in / $21.00 text outper 1M tokensabout $0.009/min blended
Google Search grounding (Gemini 3+)5,000 free requests/month, then $14per 1,000 search querieseach search query billed
Free tier$0all Live models free of charge on the free tier, rate limited; content used to improve Google products
How the per-minute estimate was worked out

Low = gemini-3.8-live, 1 min user audio (1,500 tokens x $3/1M) + 1 min model audio (1,500 tokens x $12/1M), single turn. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Live models do not support context caching, so there is no cached discount. Video input or thinking tokens add more.

Audio token rate

25 tokens per second of audio in and out = 1,500 tokens/min (Google pricing footnotes and Vertex Live billing details); 258 tokens per image.

Free tier: Yes: free tier for all Live models with lower rate limits; free-tier content may be used to improve Google products.

Source: ai.google.dev

Setup

  1. Sign in to Google AI Studio (aistudio.google.com) and create an API key for a Google Cloud project.
  2. Start on the free tier; enable billing on the project to move to paid tiers and keep data out of product improvement.
  3. Install the SDK (npm i @google/genai or pip install google-genai).
  4. Server-to-server: connect with the API key. Client-to-server: your server calls client.authTokens.create (v1beta) with uses: 1 and liveConnectConstraints, and the browser uses token.name as its apiKey.
  5. Enable contextWindowCompression and sessionResumption from day one and handle goAway by reconnecting with the latest handle.
  6. Stream 16 kHz PCM, play 24 kHz PCM, and stop playback on serverContent.interrupted.

Endpoint

wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent?key=API_KEY (ephemeral tokens use BidiGenerateContentConstrained with access_token)

Authentication

API key on the server only. Browsers and mobile use ephemeral tokens from POST https://generativelanguage.googleapis.com/v1beta/auth_tokens, sent as access_token query parameter or Authorization: Token <token>; lock model and config with liveConnectConstraints.

Quick start javascript

import { GoogleGenAI, Modality } from "@google/genai";
// Server side with an API key. In a browser, create an ephemeral token on your server
// (client.authTokens.create) and pass token.name as apiKey instead.
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const session = await ai.live.connect({
  model: "gemini-3.8-live",
  config: {
    responseModalities: [Modality.AUDIO],
    systemInstruction: "You are a friendly support agent. Keep answers short.",
    speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Kore" } } },
    contextWindowCompression: { slidingWindow: {} }, // lifts the 15 min audio session cap
    sessionResumption: {},                            // survive the ~10 min connection limit
    outputAudioTranscription: {},
  },
  callbacks: {
    onmessage: (msg) => {
      for (const part of msg.serverContent?.modelTurn?.parts ?? []) // a message can hold several parts
        if (part.inlineData?.data) playPcm24k(Buffer.from(part.inlineData.data, "base64"));
      if (msg.serverContent?.interrupted) stopPlayback(); // barge-in
      if (msg.sessionResumptionUpdate?.newHandle) saveHandle(msg.sessionResumptionUpdate.newHandle);
      if (msg.goAway) scheduleReconnect(msg.goAway.timeLeft);
    },
    onerror: (e) => console.error(e),
    onclose: (e) => console.log("closed", e.reason),
  },
});
// 16 kHz mono PCM16 chunks from the microphone
export const sendAudio = (pcm16) =>
  session.sendRealtimeInput({ audio: { data: pcm16.toString("base64"), mimeType: "audio/pcm;rate=16000" } });

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

10 minute connections, 15 minute sessions

Connections drop around 10 minutes and audio-only sessions stop at 15 minutes (2 minutes with video) unless you enable context window compression and session resumption. Every production app needs reconnect logic driven by goAway.

Whole context re-billed every turn, no caching

Google bills all tokens in the session context window on each turn and Live models do not support caching. At 25 tokens/s, long calls grow fast; use a sliding window with a sensible trigger_tokens.

Free tier trains on your data

Free-tier content is used to improve Google products. Do not run real customer calls on a free-tier key.

Model churn

Google moved from 2.0 Flash Live and Live 2.5 Flash (shut down Dec 9, 2025) to 2.5 native audio, 3.1 Flash Live preview (Mar 2026) and 3.8 Live (Sept 2026). The 2.5 models are now restricted to previous users. Config differences (thinkingLevel rejected on 3.8, async tools default) break copy-pasted setups.

Audio-only responses

Native audio models only support the AUDIO response modality. For text you must enable output audio transcription, and a single server event can carry several parts, so loop over all parts.

Manual VAD is unforgiving

With automatic VAD off the server adds no pre-speech buffer or silence tolerance; keep at least 500 ms end-of-speech silence and send audioStreamEnd when the mic pauses for over a second.

Ordering vs responsiveness

send_realtime_input favours speed over strict ordering; use send_client_content when order matters. On 3.8 models turn_complete=true in send_client_content interrupts generation.

No echo cancellation server side

Use browser echoCancellation or headphones; otherwise the model hears itself and barges in on its own speech.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Connection lifetime around 10 minutes (GoAway with timeLeft is sent before close)
  • Without compression: audio-only sessions 15 minutes, audio+video 2 minutes
  • Context window 128k tokens for native audio models (3.8 Live lists 131,072 input)
  • Session resumption tokens valid 2 hours after the last session ends (Developer API)
  • Ephemeral tokens: new-session window default 1 minute, message window default 30 minutes
  • Rate limits are per project and shown in AI Studio; Live concurrency is not published on the rate-limit page

Models and products

NameStatusNotes
gemini-3.8-liveGAGA Sept 15, 2026. Default low-latency voice agent model; 131,072 input / 65,536 output tokens; async function calling by default; thinkingLevel not supported (omit it). Supports Search grounding; no caching, structured outputs or code execution.
gemini-3.8-live-extended-thinkingGABackground reasoning (thinkingLevel low/medium/high); only NON_BLOCKING function calls; turnComplete does not mean idle, check interaction_status.
gemini-3.1-flash-live-previewPreviewReleased Mar 26, 2026; now legacy, Google recommends moving to 3.8 Live. No affective dialog or proactive audio; sequential function calling.
gemini-2.5-flash-native-audio-preview-12-2025PreviewNot deprecated but since Sept 18, 2026 limited to projects that used it before.
gemini-3.5-live-translate-previewPreviewLive speech translation.
gemini-3.5-transcribe-livePreviewStreaming STT over the Live API (Aug 26, 2026).
gemini-2.0-flash-live-001, gemini-live-2.5-flash-previewDeprecatedShut down December 9, 2025.

Docs and sources

Docs

Sources used

Not fully verified

Live API concurrent session limits per tier on the Developer API; exact language count (99 vs 70 on different pages); whether 'half-cascade' models still exist (current docs only describe native audio); latency figures. The live guide still says 'the Live API is in preview' while the models are labelled stable/GA.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?