Preview Google

Gemini Live API (video input)

Pointer entry: the Gemini Live API accepts camera or screen frames alongside audio. This entry covers only what video frames cost; see the voice segment for the rest.

Est. per minute$0.002 - 0.0042
1 high-severity warning

Overview

Best for: Cheapest realtime 'see and talk' assistant; video input adds only fractions of a cent per minute.

At a glance

Session $/min$0.002
Entry plan $/mo$0
FPS1
Max session min2
WebSocketYes
Self-hostNo
Open weightsNo

Price is video input only on Gemini 3.8 Live (audio billed on top). Audio plus video sessions capped at 2 min without session management. fps is max input frame rate.

Audio in

PCM audio

Audio out

Native audio

Languages

See voice segment.

Latency

See voice segment.

Regions

See voice segment.

Compliance

Google Cloud / Gemini API terms; see voice segment.

Features

  • camera and screen frames as JPEG/PNG
  • mediaResolution control of tokens per frame
  • session resumption / compression for longer sessions

Pricing

WhatPriceUnitNotes
Video/image input, Gemini 3.8 Live$1.00 per 1M tokens or $0.002per minutePaid tier
Tokens per video frame (Gemini 3)70 tokens (low/medium/default), 280 (high)per frameMax 1 frame per second in Live API
Audio for comparison, 3.8 Live$0.005 in / $0.018 outper minute
How the per-minute estimate was worked out

Video input only. Low = Google's listed $0.002/min; high = our arithmetic at 1 fps x 70 tokens x 60 s = 4,200 tokens at $1/M. Audio is billed on top.

Free tier: See Gemini API free tier (not checked here).

Source: ai.google.dev

Setup

  1. Get a Gemini API key (or use ephemeral tokens in browsers).
  2. Open a Live session with response_modalities AUDIO and media_resolution LOW.
  3. Send JPEG frames at 1 fps or less with send_realtime_input(video=...).
  4. Enable context window compression for sessions over 2 minutes.

Endpoint

Live API WebSocket via google-genai SDK

Authentication

GEMINI_API_KEY or ephemeral token

Quick start python

import asyncio
from google import genai
from google.genai import types

client = genai.Client()  # GEMINI_API_KEY
config = {
    "response_modalities": ["AUDIO"],
    "media_resolution": "MEDIA_RESOLUTION_LOW",  # 70 tokens/frame on Gemini 3
}

async def main(jpeg_frames):
    async with client.aio.live.connect(model="gemini-3.8-live", config=config) as s:
        for jpg in jpeg_frames:            # at most 1 frame per second
            await s.send_realtime_input(
                video=types.Blob(data=jpg, mime_type="image/jpeg"))
            await asyncio.sleep(1)
        await s.send_realtime_input(text="What do you see?")
        async for msg in s.receive():
            pass  # handle audio chunks / transcripts

# asyncio.run(main(frames))

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

2 minute cap with video

Audio plus video sessions are limited to 2 minutes unless you configure session management; plan compression or resumption.

1 fps maximum

The model sees at most one frame per second; it will miss fast motion. Not suitable for frame-accurate tasks.

Set media resolution low

Video frames default to 70 tokens on Gemini 3 but high is 280; images at high or ultra_high cost far more.

Model IDs churn

Live models move fast (3.1 Flash Live preview, 3.8 Live, 2.5 native audio preview); pin a model and watch deprecations.

Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.

Limits

  • Max 1 frame per second
  • Audio plus video sessions limited to 2 minutes unless you use session management/compression
  • Context window 128k tokens for native audio models

Models and products

NameStatusNotes
gemini-3.8-liveGAImage/video input $1.00 per 1M tokens or $0.002/min.
gemini-3.1-flash-live-previewPreviewSame price card as 3.8 Live.
gemini-2.5-flash-native-audio-preview-12-2025PreviewAudio/video input $3.00 per 1M tokens.
gemini-robotics-er-2-streaming-previewPreview$1.00/M input incl. video through 31 Dec 2026, $2.00 from 1 Jan 2027.

Docs and sources

Docs

Sources used

Not fully verified

Whether gemini-3.8-live is GA or preview label; per-frame tokens for 2.5 models.

Similar avatars + video APIs

Spotted a wrong price or a dead link?