Gemini Live API (video input)
Pointer entry: the Gemini Live API accepts camera or screen frames alongside audio. This entry covers only what video frames cost; see the voice segment for the rest.
Overview
Best for: Cheapest realtime 'see and talk' assistant; video input adds only fractions of a cent per minute.
At a glance
Price is video input only on Gemini 3.8 Live (audio billed on top). Audio plus video sessions capped at 2 min without session management. fps is max input frame rate.
PCM audio
Native audio
See voice segment.
See voice segment.
See voice segment.
Google Cloud / Gemini API terms; see voice segment.
Features
- camera and screen frames as JPEG/PNG
- mediaResolution control of tokens per frame
- session resumption / compression for longer sessions
Pricing
| What | Price | Unit |
|---|---|---|
| Video/image input, Gemini 3.8 Live | $1.00 per 1M tokens or $0.002 | per minute |
| Tokens per video frame (Gemini 3) | 70 tokens (low/medium/default), 280 (high) | per frame |
| Audio for comparison, 3.8 Live | $0.005 in / $0.018 out | per minute |
Video input only. Low = Google's listed $0.002/min; high = our arithmetic at 1 fps x 70 tokens x 60 s = 4,200 tokens at $1/M. Audio is billed on top.
Free tier: See Gemini API free tier (not checked here).
Source: ai.google.dev
Setup
- Get a Gemini API key (or use ephemeral tokens in browsers).
- Open a Live session with response_modalities AUDIO and media_resolution LOW.
- Send JPEG frames at 1 fps or less with send_realtime_input(video=...).
- Enable context window compression for sessions over 2 minutes.
Endpoint
Live API WebSocket via google-genai SDK
Authentication
GEMINI_API_KEY or ephemeral token
Quick start python
import asyncio
from google import genai
from google.genai import types
client = genai.Client() # GEMINI_API_KEY
config = {
"response_modalities": ["AUDIO"],
"media_resolution": "MEDIA_RESOLUTION_LOW", # 70 tokens/frame on Gemini 3
}
async def main(jpeg_frames):
async with client.aio.live.connect(model="gemini-3.8-live", config=config) as s:
for jpg in jpeg_frames: # at most 1 frame per second
await s.send_realtime_input(
video=types.Blob(data=jpg, mime_type="image/jpeg"))
await asyncio.sleep(1)
await s.send_realtime_input(text="What do you see?")
async for msg in s.receive():
pass # handle audio chunks / transcripts
# asyncio.run(main(frames))
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
2 minute cap with video
Audio plus video sessions are limited to 2 minutes unless you configure session management; plan compression or resumption.
1 fps maximum
The model sees at most one frame per second; it will miss fast motion. Not suitable for frame-accurate tasks.
Set media resolution low
Video frames default to 70 tokens on Gemini 3 but high is 280; images at high or ultra_high cost far more.
Model IDs churn
Live models move fast (3.1 Flash Live preview, 3.8 Live, 2.5 native audio preview); pin a model and watch deprecations.
Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.
Limits
- Max 1 frame per second
- Audio plus video sessions limited to 2 minutes unless you use session management/compression
- Context window 128k tokens for native audio models
Models and products
| Name | Status |
|---|---|
| gemini-3.8-live | GA |
| gemini-3.1-flash-live-preview | Preview |
| gemini-2.5-flash-native-audio-preview-12-2025 | Preview |
| gemini-robotics-er-2-streaming-preview | Preview |
Docs and sources
Docs
Sources used
- ai.google.dev/gemini-api/docs/pricing
- ai.google.dev/gemini-api/docs/media-resolution
- ai.google.dev/gemini-api/docs/live-api/capabilities
Whether gemini-3.8-live is GA or preview label; per-frame tokens for 2.5 models.