OpenAI GPT-Live API
Full-duplex voice model (gpt-live-1) that listens while it speaks and hands reasoning and tools to a backend agent, billed per minute. Same voice system as ChatGPT Voice; suits natural, interruption-heavy conversations where you already have an agent backend.
Overview
Best for: Natural, full-duplex conversations (companions, concierge, phone agents) where an existing agent or Responses model does the thinking and you want simple per-minute voice pricing.
At a glance
Concurrency is the Build tier (Launch 300, Grow 500). $0.05/min covers the voice layer only; delegated backend model and tools billed separately. Can delegate to your own agent backend. Session has an unstated duration limit; outbound SIP calls cap at 2 hours.
WebSocket: audio/pcm 24 kHz (default) or 16 kHz mono 16-bit LE, audio/pcmu or audio/pcma 8 kHz; raw bytes base64, even byte length, no container. Format fixed at session start. WebRTC and SIP negotiate codecs (outbound SIP needs Opus + SDES-SRTP).
Same format as input (one setting covers both).
Not listed on the model page.
12 new voices listed in the docs (quartz, ripple, vesper, willow, stone, gleam, meridian, bossa, tempo, beacon, delta, cinder); docs give marin as the default. Voice change requires a new session. Audio is SynthID watermarked (reported July 31, 2026 update).
Vendor claim: improves Full Duplex Bench score by 30 percentage points over gpt-realtime-2.1 (reported via third-party coverage). No millisecond figure published.
Global API; data residency in the United States and Europe for gpt-live-1. Also offered on Azure (see Azure entry) with its own concurrent-session tiers.
/v1/live/sessions is ZDR eligible with limitations (store forced false, no forking or recording download). Abuse-monitoring logs 30 days. US and EU data residency.
Features
- full duplex (listens while speaking), smooth interruption handling
- delegation to a Responses backend managed by OpenAI, or client delegation to your own agent
- function calling and backend tools (e.g. web_search)
- input and output transcript deltas
- context appends during the session (up to 500 tokens per event)
- startup history seeding (up to 128 messages / 8,192 tokens)
- session forking from a stored recording (not under ZDR)
- sideband server connection to monitor or control a session
- inbound and outbound SIP calling (outbound must be enabled for your org)
- partner integrations: LiveKit, Twilio, Telnyx, Daily/Pipecat
Pricing
| What | Price | Unit |
|---|---|---|
| gpt-live-1 voice session | $0.05 | per minute, billed per second (not rounded up) |
| Backend model and tools (Responses delegation) | standard model rates | per token / per tool call |
Low = voice layer only, $0.05 per minute of session duration with no delegated work. High depends entirely on the backend model, how often the model delegates and tool fees; budget per-minute session cost plus backend token spend measured in your own tests.
Not applicable for the voice layer (duration billing). Backend usage is ordinary text tokens.
Free tier: Free tier not supported.
Source: developers.openai.com
Setup
- Use a paid OpenAI project (free tier unsupported) and create a project API key kept on your server.
- Browser: create an RTCPeerConnection with the mic track and an oai-events data channel, send the SDP offer to your own server.
- Server: POST /v1/live/sessions with session.model = gpt-live-1, delegation settings and transport { type: webrtc, sdp }; return the SDP answer from the 201 response.
- Wait for session.started on the data channel before speaking; do not send session.start over WebRTC.
- Server-side audio: open wss://api.openai.com/v1/live/sessions and send session.start as the first message.
- End with session.close and wait for session.closed to collect final usage.
Endpoint
POST https://api.openai.com/v1/live/sessions (WebRTC session create); wss://api.openai.com/v1/live/sessions (WebSocket); SIP via live.transport.incoming webhook and POST /v1/live/sessions/{session_id}/accept
Authentication
Project API key on your server. The docs do not document an ephemeral client-secret flow for Live: keep the browser talking to your server for the SDP exchange.
Quick start javascript
import WebSocket from "ws"; // server side; browsers should use WebRTC via your server
const ws = new WebSocket("wss://api.openai.com/v1/live/sessions", {
headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },
});
ws.on("open", () => ws.send(JSON.stringify({
type: "session.start",
session: {
model: "gpt-live-1",
instructions: "Be concise. Delegate anything that needs current information.",
audio: { format: { type: "audio/pcm", rate: 24000 }, output: { voice: "marin" } },
delegation: {
type: "responses",
responses: { model: "gpt-5.6-luna", tools: [{ type: "web_search" }], tool_choice: "auto" },
},
},
})));
let ready = false;
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "session.started") ready = true; // do not send audio before this
if (ev.type === "session.output_audio.delta") playPcm16(Buffer.from(ev.delta, "base64"));
if (ev.type === "session.closed") console.log("final usage", ev);
if (ev.type === "error") console.error(ev.error);
});
// 24 kHz mono PCM16, even byte length, no WAV header
export function sendAudio(pcm16) {
if (ready) ws.send(JSON.stringify({ type: "session.input_audio.append", audio: pcm16.toString("base64") }));
}
// To end: ws.send(JSON.stringify({ type: "session.close" })) and wait for "session.closed"
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Two bills, not one
$0.05/min covers only the voice layer. Every delegated task runs a Responses model plus tools billed at normal rates, which can easily exceed the voice cost on tool-heavy calls. Log backend usage per session.
Billed by wall-clock duration
Billing is per second of session duration, so idle or forgotten sessions keep costing. Always send session.close and set your own inactivity timeout.
Long calls lose memory
When context passes 90 percent of 128k, GPT-Live swaps to a replacement engine with at most 8,192 tokens of history. Keep critical facts in instructions or re-append them.
Different API from Realtime
GPT-Live uses /v1/live/sessions with session.start, session.input_audio.append and session.output_audio.delta. Realtime code and events do not carry over; follow the migration guide. Delegation mode cannot be changed mid-session.
Transcript events need stitching
Transcript deltas have no item id or turn-complete event and output audio deltas have no timing fields, so your app must group them itself.
Moderation can cut audio
Moderation can stop assistant audio without ending the session, or end the session. Handle error events with client_event_id and show the user something.
Outbound calls are not idempotent
Each outbound create places a new call and X-Client-Request-Id does not deduplicate. Never auto-retry after an ambiguous timeout.
Recordings and ZDR
store: true keeps recordings 30 days; under ZDR store is forced false and forking or recording download is unavailable.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Concurrent sessions (OpenAI direct): Build 50, Launch 300, Grow 500; Free unsupported
- Context window 128,000 tokens; above 90 percent usage a replacement engine starts with up to 8,192 tokens of history, so older details may be summarised or dropped
- Instructions up to 16,384 tokens; startup history up to 128 messages / 8,192 tokens
- Session ends with reason expired at a duration limit that the docs do not state
- Outbound SIP: 3 minutes ringing, 2 hours max connected call, 1 MiB request body (not configurable)
- Stored recordings kept 30 days when store: true
Models and products
| Name | Status |
|---|---|
| gpt-live-1 | GA |
Docs and sources
Docs
Sources used
- developers.openai.com/api/docs/models/gpt-live-1
- developers.openai.com/api/docs/guides/live
- developers.openai.com/api/docs/guides/live-conversations
- developers.openai.com/api/docs/guides/voice-webrtc?api=live
- developers.openai.com/api/docs/guides/voice-websockets?api=live
- developers.openai.com/api/docs/guides/voice-sip?api=realtime
- developers.openai.com/api/docs/guides/your-data
- windowsreport.com/gpt-live-1-is-now-available-to-developers-through-the-openai-...
- unite.ai/openais-gpt-live-1-arrives-in-the-api-at-0-05-per-minute
Formal GA/preview label (model page silent); maximum session duration; supported languages; voice list vs. default (docs list 12 names but give marin as default); launch date (Sept 10, 2026 from secondary sources); latency figures. The Free-tier/concurrency table differs slightly between search snippets (Tier 1-5) and the current model page (Build/Launch/Grow).