OpenAI Realtime API
Native speech-to-speech API for the gpt-realtime family over WebRTC, WebSocket or SIP, plus realtime transcription and live translation sessions. The default choice for production voice agents that need strong tool calling and reasoning.
Overview
Best for: Production voice agents and phone bots that need the strongest tool calling and reasoning with a mature SDK, WebRTC in the browser and native SIP.
At a glance
Prices are gpt-realtime-2.1 (mini: $10/$20). SOC 2 and BAA are account-level OpenAI programs; confirm realtime coverage. EU residency needs approved abuse-monitoring controls. Languages: multilingual, no published list. Latency: only a relative p95 claim.
audio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk.
audio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response.
Multilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents.
10 built-in voices: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar. OpenAI recommends marin or cedar. No custom voices.
Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published.
Global API. Data residency in the United States and Europe (EEA + Switzerland) for gpt-realtime, -1.5, -mini, -2, -2.1, -2.1-mini; EU requires approved abuse-monitoring controls (ZDR, Modified Abuse Monitoring, etc). EU SIP endpoint sip-eu.api.openai.com. Tracing is not EU-residency compliant for /v1/realtime.
/v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement.
Features
- function calling (parallel tool calls)
- remote MCP servers as tools
- semantic_vad and server_vad turn detection, or manual push-to-talk (turn_detection null)
- barge-in with conversation.item.truncate to sync what the user heard
- image input (input_image) on gpt-realtime-2.x and gpt-realtime
- configurable reasoning effort on 2.x models
- input audio transcription in parallel
- out-of-band responses (response.conversation = none)
- automatic prompt caching
- truncation controls (retention_ratio, token_limits.post_instructions)
- SIP telephony with webhooks (realtime.call.incoming)
- sideband server WebSocket for a WebRTC or SIP call (?call_id=)
- realtime transcription sessions and live translation sessions
- Agents SDK (@openai/agents/realtime) for browser and server
Pricing
| What | Price | Unit |
|---|---|---|
| gpt-realtime-2.1 audio input | $32.00 | per 1M tokens |
| gpt-realtime-2.1 audio output | $64.00 | per 1M tokens |
| gpt-realtime-2.1 text input | $4.00 | per 1M tokens |
| gpt-realtime-2.1 text output | $24.00 | per 1M tokens |
| gpt-realtime-2.1 image input | $5.00 | per 1M tokens |
| gpt-realtime-2.1-mini audio input | $10.00 | per 1M tokens |
| gpt-realtime-2.1-mini audio output | $20.00 | per 1M tokens |
| gpt-realtime-2.1-mini text input / output | $0.60 / $2.40 | per 1M tokens |
| gpt-realtime-2.1-mini image input | $0.80 | per 1M tokens |
| gpt-realtime-2 (all lines) | same as gpt-realtime-2.1 | per 1M tokens |
| gpt-realtime-1.5 / gpt-realtime audio in / out | $32.00 / $64.00 | per 1M tokens |
| gpt-realtime-mini audio in / out | $10.00 / $20.00 | per 1M tokens |
| gpt-realtime-translate | $0.034 | per minute of audio |
| gpt-live-transcribe / gpt-realtime-whisper | $0.017 | per minute |
| gpt-transcribe | $0.0045 | per minute |
| gpt-4o-transcribe / gpt-4o-mini-transcribe (input transcription) | $0.006 / $0.003 | per minute (estimated) |
Low = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached).
User audio 1 token per 100 ms = 10 tokens/s = 600 tokens/min. Assistant audio 1 token per 50 ms = 20 tokens/s = 1,200 tokens/min. VAD-filtered silence is not billed. (OpenAI realtime-costs guide)
Free tier: None. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow.
Source: developers.openai.com
Setup
- Create an OpenAI Platform account and a project at platform.openai.com; add a payment method (free tier cannot use realtime).
- Create a project API key and store it only on your server.
- Browser: your server calls POST https://api.openai.com/v1/realtime/client_secrets with the session config and returns the ek_ value; the browser POSTs its SDP offer to https://api.openai.com/v1/realtime/calls with that key and opens the oai-events data channel.
- Server or telephony bridge: open the WebSocket below with the API key and send session.update with session.type = realtime.
- Phone: point your SIP trunk at sip:<project_id>@sip.api.openai.com;transport=tls, add a realtime.call.incoming webhook, then POST /v1/realtime/calls/{call_id}/accept.
- Log response.done usage on every turn so you can see context growth and cache hit rates.
Endpoint
wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1 (WebSocket); https://api.openai.com/v1/realtime/calls (WebRTC SDP); sip:$PROJECT_ID@sip.api.openai.com;transport=tls (SIP)
Authentication
Authorization: Bearer <API key> for server WebSocket and SIP control. For browsers and mobile, mint an ephemeral key (ek_...) server side via POST /v1/realtime/client_secrets (expires_after.seconds 10 to 7200, default 600) and lock instructions/tools in that session config. Optional OpenAI-Safety-Identifier header on the server request is bound to the token.
Quick start javascript
import WebSocket from "ws"; // server side only: never ship the API key to a browser
const ws = new WebSocket("wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1", {
headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },
});
ws.on("open", () => {
ws.send(JSON.stringify({
type: "session.update",
session: {
type: "realtime",
instructions: "You are a friendly support agent. Keep answers short.",
audio: {
input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: { type: "semantic_vad" } },
output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" },
},
},
}));
});
// Call for every ~100 ms chunk of 24 kHz mono PCM16 from your mic or phone bridge
export function sendAudio(pcm16) {
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.output_audio.delta") playPcm16(Buffer.from(ev.delta, "base64"));
if (ev.type === "input_audio_buffer.speech_started") stopPlayback(); // barge-in: flush local audio
if (ev.type === "response.done") console.log(ev.response.usage); // log tokens per turn
if (ev.type === "error") console.error(ev.error);
});
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Context growth multiplies cost
Each response re-sends the whole conversation, including earlier audio (model audio is re-billed as audio input at $32/1M). A 10 minute call can cost 8x the naive per-minute figure with no cache hits. Changing instructions or tools mid-session, or truncating every turn, breaks the cache. Use token_limits.post_instructions with retention_ratio around 0.8, or delete or summarise old items.
Beta interface is gone
The OpenAI-Beta: realtime=v1 interface was removed May 12, 2026. GA requires session.type, nests audio config under session.audio.input/output, and renames events (response.output_audio.delta, response.output_text.delta, response.output_audio_transcript.delta). Old tutorials and Azure preview samples using response.audio.delta will silently get no audio.
Deprecation calendar
gpt-realtime, gpt-realtime-mini, gpt-4o-realtime and gpt-4o-mini-realtime shut down January 20, 2027. gpt-4o realtime previews already died May 7, 2026. Pin explicit model ids and test 2.1 before January.
Voice is locked after first audio
Once the model has emitted audio in a session the voice cannot be changed. Set it in the client secret or the first session.update.
60 minute hard cap without warning
Sessions end at 60 minutes and no warning event is sent. Track session age yourself and reconnect at a turn boundary, replaying a summary into the new session.
Reasoning effort trades latency and cost
gpt-realtime-2.x are reasoning models; higher reasoning effort increases both time to first audio and billed text output tokens ($24/1M). Keep effort low for chit-chat and raise it only for tool-heavy turns.
Mini models are less reliable with tools
OpenAI's own cost guide warns the mini model may follow instructions and call functions less reliably. Prototype on the full model, then measure regressions before switching.
Translation endpoint has its own retention rules
/v1/realtime/translations is not listed in the ZDR column of OpenAI's data controls table, unlike /v1/realtime. Check with OpenAI before sending regulated audio.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Max session length 60 minutes, no warning event before cutoff (third-party SDKs reconnect at ~50 min)
- Context window 128,000 tokens; max output 32,000 (2.1 and 2.1-mini)
- Default rate limits gpt-realtime-2.1: Build 400 RPM / 200,000 TPM, Launch 10,000 RPM / 4,000,000 TPM, Grow 20,000 RPM / 15,000,000 TPM
- gpt-realtime-translate: Build 200, Launch 650, Grow 850 minutes of audio per minute
- Ephemeral client secret lifetime 10 s to 7,200 s, default 600 s
- Audio append chunk max 15 MB
Models and products
| Name | Status |
|---|---|
| gpt-realtime-2.1 | GA |
| gpt-realtime-2.1-mini | GA |
| gpt-realtime-2 | GA |
| gpt-realtime-1.5 | GA |
| gpt-realtime | Deprecated |
| gpt-realtime-mini | Deprecated |
| gpt-realtime-translate | GA |
| gpt-live-transcribe / gpt-realtime-whisper / gpt-transcribe | GA |
| gpt-4o-realtime-preview (all snapshots), gpt-4o-mini-realtime-preview | Deprecated |
Docs and sources
Docs
- Realtime guide
- Realtime conversations (events, voices, VAD)
- Managing costs
- WebRTC guide
- SIP guide
- Realtime transcription
- Deprecations
Sources used
- developers.openai.com/api/docs/pricing
- developers.openai.com/api/docs/models/gpt-realtime-2.1
- developers.openai.com/api/docs/models/gpt-realtime-2.1-mini
- developers.openai.com/api/docs/models/gpt-realtime-translate
- developers.openai.com/api/docs/guides/realtime-costs
- developers.openai.com/api/docs/guides/realtime-conversations.md
- developers.openai.com/api/docs/guides/voice-websockets?api=realtime
- developers.openai.com/api/docs/guides/voice-webrtc?api=realtime
- developers.openai.com/api/docs/guides/voice-sip?api=realtime
- developers.openai.com/api/reference/resources/realtime/subresources/client_secr...
- developers.openai.com/api/docs/deprecations
- developers.openai.com/api/docs/guides/your-data
- datanorth.ai/news/openai-releases-gpt-realtime-2-1-voice-models
Exact reasoning-effort levels and how to set them for gpt-realtime-2.x; latency numbers (only the third-party-reported 25 percent p95 improvement); the language lists for gpt-realtime-translate; codecs accepted on inbound Realtime SIP; whether Realtime sessions have a per-tier concurrent session cap (model page lists RPM/TPM only).