GA OpenAI

OpenAI Realtime API (image input)

Pointer entry: OpenAI's realtime speech models accept still images in the conversation. There is no continuous video track; apps send sampled frames (often 1 fps) as input_image items.

Est. per minuten/a

Overview

Best for: Voice assistants that occasionally need to look at a screenshot or camera frame.

At a glance

Entry plan $/mo$0
Free tierNo
WebRTCYes
WebSocketYes
Self-hostNo
Open weightsNo

Still images only (no video track), billed as image tokens ($5/M on gpt-realtime-2.1, $0.80/M on mini). Per-minute cost depends on frame rate and size.

Audio in

Speech

Audio out

Speech

Languages

See voice segment.

Latency

See voice segment.

Regions

See voice segment.

Compliance

OpenAI API terms; see voice segment.

Features

  • input_image items with base64 data URL
  • text + image in one message
  • frame sampling done by your app (LiveKit samples 1 fps by default)

Pricing

WhatPriceUnitNotes
Image input, gpt-realtime-2.1$5.00per 1M tokensCached $0.50
Image input, gpt-realtime-2.1-mini$0.80per 1M tokensCached $0.08
gpt-live-1 session$0.05per minuteBackend model and tool use billed separately
How the per-minute estimate was worked out

Depends on tokens per image (detail level, size) which we did not verify for realtime models

Free tier: None.

Source: developers.openai.com

Setup

  1. Open a Realtime session (WebRTC in browser with an ephemeral key, or WebSocket server-side).
  2. Capture a frame from the camera or screen to a small JPEG (e.g. 512 px).
  3. Send conversation.item.create with input_text + input_image, then response.create.
  4. Throttle frames; only send on change or on user request.

Endpoint

wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1

Authentication

Bearer API key server-side; ephemeral client secret in browser

Quick start javascript

// ws: an open Realtime WebSocket (or RTCDataChannel "oai-events")
function sendFrame(ws, base64Jpeg, question) {
  ws.send(JSON.stringify({
    type: "conversation.item.create",
    item: {
      type: "message",
      role: "user",
      content: [
        { type: "input_text", text: question },
        { type: "input_image", image_url: `data:image/jpeg;base64,${base64Jpeg}` },
      ],
    },
  }));
  ws.send(JSON.stringify({ type: "response.create" }));
}

// Grab a frame: draw <video> to a 512px canvas, then
// canvas.toDataURL("image/jpeg", 0.7).split(",")[1]

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Images stay in context and cost every turn

Each image becomes conversation tokens that are re-read on later turns; prune old frames or you pay for them repeatedly (cached rate helps).

Not real video understanding

Sampled stills miss motion; for continuous monitoring use a dedicated video service.

Prefer WebSocket for big images

Developers reported trouble sending large base64 images over the WebRTC data channel; downscale to around 512 px.

gpt-live-1 vision unclear

We could not confirm that gpt-live-1 accepts images; test before choosing it for vision.

Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.

Limits

  • No native video stream; you choose frame rate and size
  • Large base64 images over WebRTC data channel reported as troublesome by developers

Models and products

NameStatusNotes
gpt-realtime-2.1GAImage input $5.00 / cached $0.50 per 1M tokens.
gpt-realtime-2.1-miniGAImage input $0.80 / cached $0.08 per 1M tokens.
gpt-realtime-2, gpt-realtime-1.5, gpt-realtimeGASame price card as 2.1.
gpt-realtime-miniGASame as 2.1-mini.
gpt-live-1GAListed as GPT-Live sessions at $0.05/min; image/video support not confirmed.

Docs and sources

Docs

Sources used

Not fully verified

Tokens per image for realtime models, gpt-live-1 vision support.

Similar avatars + video APIs

Spotted a wrong price or a dead link?