GA OpenAI

OpenAI GPT-Live API

Full-duplex voice model (gpt-live-1) that listens while it speaks and hands reasoning and tools to a backend agent, billed per minute. Same voice system as ChatGPT Voice; suits natural, interruption-heavy conversations where you already have an agent backend.

Est. per minute$0.05
2 high-severity warnings

Overview

Best for: Natural, full-duplex conversations (companions, concierge, phone agents) where an existing agent or Responses model does the thinking and you want simple per-minute voice pricing.

At a glance

Flat $/min$0.05
Free tierNo
Native S2SYes
ToolsYes
Image inNo
Own LLMYes
Voices12
Context tokens128,000
Concurrency50
WebRTCYes
WebSocketYes
Phone / SIPYes
EU dataYes
Open weightsNo

Concurrency is the Build tier (Launch 300, Grow 500). $0.05/min covers the voice layer only; delegated backend model and tools billed separately. Can delegate to your own agent backend. Session has an unstated duration limit; outbound SIP calls cap at 2 hours.

Audio in

WebSocket: audio/pcm 24 kHz (default) or 16 kHz mono 16-bit LE, audio/pcmu or audio/pcma 8 kHz; raw bytes base64, even byte length, no container. Format fixed at session start. WebRTC and SIP negotiate codecs (outbound SIP needs Opus + SDES-SRTP).

Audio out

Same format as input (one setting covers both).

Languages

Not listed on the model page.

Voices

12 new voices listed in the docs (quartz, ripple, vesper, willow, stone, gleam, meridian, bossa, tempo, beacon, delta, cinder); docs give marin as the default. Voice change requires a new session. Audio is SynthID watermarked (reported July 31, 2026 update).

Latency

Vendor claim: improves Full Duplex Bench score by 30 percentage points over gpt-realtime-2.1 (reported via third-party coverage). No millisecond figure published.

Regions

Global API; data residency in the United States and Europe for gpt-live-1. Also offered on Azure (see Azure entry) with its own concurrent-session tiers.

Compliance

/v1/live/sessions is ZDR eligible with limitations (store forced false, no forking or recording download). Abuse-monitoring logs 30 days. US and EU data residency.

Features

  • full duplex (listens while speaking), smooth interruption handling
  • delegation to a Responses backend managed by OpenAI, or client delegation to your own agent
  • function calling and backend tools (e.g. web_search)
  • input and output transcript deltas
  • context appends during the session (up to 500 tokens per event)
  • startup history seeding (up to 128 messages / 8,192 tokens)
  • session forking from a stored recording (not under ZDR)
  • sideband server connection to monitor or control a session
  • inbound and outbound SIP calling (outbound must be enabled for your org)
  • partner integrations: LiveKit, Twilio, Telnyx, Daily/Pipecat

Pricing

WhatPriceUnitNotes
gpt-live-1 voice session$0.05per minute, billed per second (not rounded up)covers the voice layer only
Backend model and tools (Responses delegation)standard model ratesper token / per tool callbilled separately at the normal price of the chosen backend model and tools
How the per-minute estimate was worked out

Low = voice layer only, $0.05 per minute of session duration with no delegated work. High depends entirely on the backend model, how often the model delegates and tool fees; budget per-minute session cost plus backend token spend measured in your own tests.

Audio token rate

Not applicable for the voice layer (duration billing). Backend usage is ordinary text tokens.

Free tier: Free tier not supported.

Source: developers.openai.com

Setup

  1. Use a paid OpenAI project (free tier unsupported) and create a project API key kept on your server.
  2. Browser: create an RTCPeerConnection with the mic track and an oai-events data channel, send the SDP offer to your own server.
  3. Server: POST /v1/live/sessions with session.model = gpt-live-1, delegation settings and transport { type: webrtc, sdp }; return the SDP answer from the 201 response.
  4. Wait for session.started on the data channel before speaking; do not send session.start over WebRTC.
  5. Server-side audio: open wss://api.openai.com/v1/live/sessions and send session.start as the first message.
  6. End with session.close and wait for session.closed to collect final usage.

Endpoint

POST https://api.openai.com/v1/live/sessions (WebRTC session create); wss://api.openai.com/v1/live/sessions (WebSocket); SIP via live.transport.incoming webhook and POST /v1/live/sessions/{session_id}/accept

Authentication

Project API key on your server. The docs do not document an ephemeral client-secret flow for Live: keep the browser talking to your server for the SDP exchange.

Quick start javascript

import WebSocket from "ws"; // server side; browsers should use WebRTC via your server
const ws = new WebSocket("wss://api.openai.com/v1/live/sessions", {
  headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },
});
ws.on("open", () => ws.send(JSON.stringify({
  type: "session.start",
  session: {
    model: "gpt-live-1",
    instructions: "Be concise. Delegate anything that needs current information.",
    audio: { format: { type: "audio/pcm", rate: 24000 }, output: { voice: "marin" } },
    delegation: {
      type: "responses",
      responses: { model: "gpt-5.6-luna", tools: [{ type: "web_search" }], tool_choice: "auto" },
    },
  },
})));
let ready = false;
ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "session.started") ready = true; // do not send audio before this
  if (ev.type === "session.output_audio.delta") playPcm16(Buffer.from(ev.delta, "base64"));
  if (ev.type === "session.closed") console.log("final usage", ev);
  if (ev.type === "error") console.error(ev.error);
});
// 24 kHz mono PCM16, even byte length, no WAV header
export function sendAudio(pcm16) {
  if (ready) ws.send(JSON.stringify({ type: "session.input_audio.append", audio: pcm16.toString("base64") }));
}
// To end: ws.send(JSON.stringify({ type: "session.close" })) and wait for "session.closed"

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Two bills, not one

$0.05/min covers only the voice layer. Every delegated task runs a Responses model plus tools billed at normal rates, which can easily exceed the voice cost on tool-heavy calls. Log backend usage per session.

Billed by wall-clock duration

Billing is per second of session duration, so idle or forgotten sessions keep costing. Always send session.close and set your own inactivity timeout.

Long calls lose memory

When context passes 90 percent of 128k, GPT-Live swaps to a replacement engine with at most 8,192 tokens of history. Keep critical facts in instructions or re-append them.

Different API from Realtime

GPT-Live uses /v1/live/sessions with session.start, session.input_audio.append and session.output_audio.delta. Realtime code and events do not carry over; follow the migration guide. Delegation mode cannot be changed mid-session.

Transcript events need stitching

Transcript deltas have no item id or turn-complete event and output audio deltas have no timing fields, so your app must group them itself.

Moderation can cut audio

Moderation can stop assistant audio without ending the session, or end the session. Handle error events with client_event_id and show the user something.

Outbound calls are not idempotent

Each outbound create places a new call and X-Client-Request-Id does not deduplicate. Never auto-retry after an ambiguous timeout.

Recordings and ZDR

store: true keeps recordings 30 days; under ZDR store is forced false and forking or recording download is unavailable.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Concurrent sessions (OpenAI direct): Build 50, Launch 300, Grow 500; Free unsupported
  • Context window 128,000 tokens; above 90 percent usage a replacement engine starts with up to 8,192 tokens of history, so older details may be summarised or dropped
  • Instructions up to 16,384 tokens; startup history up to 128 messages / 8,192 tokens
  • Session ends with reason expired at a duration limit that the docs do not state
  • Outbound SIP: 3 minutes ringing, 2 hours max connected call, 1 MiB request body (not configurable)
  • Stored recordings kept 30 days when store: true

Models and products

NameStatusNotes
gpt-live-1GAAPI launch reported Sept 10, 2026 (third-party coverage); model page does not label preview or GA. Audio and text in and out, no image or video. Knowledge cutoff Jul 31, 2025. Delegates to a Responses model (e.g. gpt-5.6-terra or gpt-5.6-luna) or to your own backend.

Docs and sources

Docs

Sources used

Not fully verified

Formal GA/preview label (model page silent); maximum session duration; supported languages; voice list vs. default (docs list 12 names but give marin as default); launch date (Sept 10, 2026 from secondary sources); latency figures. The Free-tier/concurrency table differs slightly between search snippets (Tier 1-5) and the current model page (Build/Launch/Grow).

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?