OpenAI Realtime API (image input)
Pointer entry: OpenAI's realtime speech models accept still images in the conversation. There is no continuous video track; apps send sampled frames (often 1 fps) as input_image items.
Overview
Best for: Voice assistants that occasionally need to look at a screenshot or camera frame.
At a glance
Still images only (no video track), billed as image tokens ($5/M on gpt-realtime-2.1, $0.80/M on mini). Per-minute cost depends on frame rate and size.
Speech
Speech
See voice segment.
See voice segment.
See voice segment.
OpenAI API terms; see voice segment.
Features
- input_image items with base64 data URL
- text + image in one message
- frame sampling done by your app (LiveKit samples 1 fps by default)
Pricing
| What | Price | Unit |
|---|---|---|
| Image input, gpt-realtime-2.1 | $5.00 | per 1M tokens |
| Image input, gpt-realtime-2.1-mini | $0.80 | per 1M tokens |
| gpt-live-1 session | $0.05 | per minute |
Depends on tokens per image (detail level, size) which we did not verify for realtime models
Free tier: None.
Source: developers.openai.com
Setup
- Open a Realtime session (WebRTC in browser with an ephemeral key, or WebSocket server-side).
- Capture a frame from the camera or screen to a small JPEG (e.g. 512 px).
- Send conversation.item.create with input_text + input_image, then response.create.
- Throttle frames; only send on change or on user request.
Endpoint
wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1
Authentication
Bearer API key server-side; ephemeral client secret in browser
Quick start javascript
// ws: an open Realtime WebSocket (or RTCDataChannel "oai-events")
function sendFrame(ws, base64Jpeg, question) {
ws.send(JSON.stringify({
type: "conversation.item.create",
item: {
type: "message",
role: "user",
content: [
{ type: "input_text", text: question },
{ type: "input_image", image_url: `data:image/jpeg;base64,${base64Jpeg}` },
],
},
}));
ws.send(JSON.stringify({ type: "response.create" }));
}
// Grab a frame: draw <video> to a 512px canvas, then
// canvas.toDataURL("image/jpeg", 0.7).split(",")[1]
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Images stay in context and cost every turn
Each image becomes conversation tokens that are re-read on later turns; prune old frames or you pay for them repeatedly (cached rate helps).
Not real video understanding
Sampled stills miss motion; for continuous monitoring use a dedicated video service.
Prefer WebSocket for big images
Developers reported trouble sending large base64 images over the WebRTC data channel; downscale to around 512 px.
gpt-live-1 vision unclear
We could not confirm that gpt-live-1 accepts images; test before choosing it for vision.
Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.
Limits
- No native video stream; you choose frame rate and size
- Large base64 images over WebRTC data channel reported as troublesome by developers
Models and products
| Name | Status |
|---|---|
| gpt-realtime-2.1 | GA |
| gpt-realtime-2.1-mini | GA |
| gpt-realtime-2, gpt-realtime-1.5, gpt-realtime | GA |
| gpt-realtime-mini | GA |
| gpt-live-1 | GA |
Docs and sources
Docs
Sources used
- developers.openai.com/api/docs/pricing
- community.openai.com/t/realtime-model-image-input/1355688
- docs.livekit.io/agents/models/realtime/openai/
Tokens per image for realtime models, gpt-live-1 vision support.