GA OpenAI

OpenAI Realtime API

Native speech-to-speech API for the gpt-realtime family over WebRTC, WebSocket or SIP, plus realtime transcription and live translation sessions. The default choice for production voice agents that need strong tool calling and reasoning.

Est. per minute$0.096 - 0.76
3 high-severity warnings

Overview

Best for: Production voice agents and phone bots that need the strongest tool calling and reasoning with a mature SDK, WebRTC in the browser and native SIP.

At a glance

Audio in $/1M tok$32
Audio out $/1M tok$64
Free tierNo
Native S2SYes
ToolsYes
Image inYes
Own LLMNo
Voices10
CloningNo
Context tokens128,000
Max session min60
WebRTCYes
WebSocketYes
Phone / SIPYes
HIPAAYes
SOC 2Yes
EU dataYes
Open weightsNo

Prices are gpt-realtime-2.1 (mini: $10/$20). SOC 2 and BAA are account-level OpenAI programs; confirm realtime coverage. EU residency needs approved abuse-monitoring controls. Languages: multilingual, no published list. Latency: only a relative p95 claim.

Audio in

audio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk.

Audio out

audio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response.

Languages

Multilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents.

Voices

10 built-in voices: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar. OpenAI recommends marin or cedar. No custom voices.

Latency

Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published.

Regions

Global API. Data residency in the United States and Europe (EEA + Switzerland) for gpt-realtime, -1.5, -mini, -2, -2.1, -2.1-mini; EU requires approved abuse-monitoring controls (ZDR, Modified Abuse Monitoring, etc). EU SIP endpoint sip-eu.api.openai.com. Tracing is not EU-residency compliant for /v1/realtime.

Compliance

/v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement.

Features

  • function calling (parallel tool calls)
  • remote MCP servers as tools
  • semantic_vad and server_vad turn detection, or manual push-to-talk (turn_detection null)
  • barge-in with conversation.item.truncate to sync what the user heard
  • image input (input_image) on gpt-realtime-2.x and gpt-realtime
  • configurable reasoning effort on 2.x models
  • input audio transcription in parallel
  • out-of-band responses (response.conversation = none)
  • automatic prompt caching
  • truncation controls (retention_ratio, token_limits.post_instructions)
  • SIP telephony with webhooks (realtime.call.incoming)
  • sideband server WebSocket for a WebRTC or SIP call (?call_id=)
  • realtime transcription sessions and live translation sessions
  • Agents SDK (@openai/agents/realtime) for browser and server

Pricing

WhatPriceUnitNotes
gpt-realtime-2.1 audio input$32.00per 1M tokenscached $0.40
gpt-realtime-2.1 audio output$64.00per 1M tokens
gpt-realtime-2.1 text input$4.00per 1M tokenscached $0.40
gpt-realtime-2.1 text output$24.00per 1M tokensincludes reasoning tokens; higher reasoning effort raises output usage
gpt-realtime-2.1 image input$5.00per 1M tokenscached $0.50
gpt-realtime-2.1-mini audio input$10.00per 1M tokenscached $0.30
gpt-realtime-2.1-mini audio output$20.00per 1M tokens
gpt-realtime-2.1-mini text input / output$0.60 / $2.40per 1M tokenscached input $0.06
gpt-realtime-2.1-mini image input$0.80per 1M tokenscached $0.08
gpt-realtime-2 (all lines)same as gpt-realtime-2.1per 1M tokens
gpt-realtime-1.5 / gpt-realtime audio in / out$32.00 / $64.00per 1M tokenstext $4.00 in, $16.00 out; cached $0.40
gpt-realtime-mini audio in / out$10.00 / $20.00per 1M tokenstext $0.60 / $2.40
gpt-realtime-translate$0.034per minute of audiobilled by duration
gpt-live-transcribe / gpt-realtime-whisper$0.017per minuterealtime transcription sessions
gpt-transcribe$0.0045per minute
gpt-4o-transcribe / gpt-4o-mini-transcribe (input transcription)$0.006 / $0.003per minute (estimated)token rates $2.50/$10.00 and $1.25/$5.00 per 1M
How the per-minute estimate was worked out

Low = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached).

Audio token rate

User audio 1 token per 100 ms = 10 tokens/s = 600 tokens/min. Assistant audio 1 token per 50 ms = 20 tokens/s = 1,200 tokens/min. VAD-filtered silence is not billed. (OpenAI realtime-costs guide)

Free tier: None. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow.

Source: developers.openai.com

Setup

  1. Create an OpenAI Platform account and a project at platform.openai.com; add a payment method (free tier cannot use realtime).
  2. Create a project API key and store it only on your server.
  3. Browser: your server calls POST https://api.openai.com/v1/realtime/client_secrets with the session config and returns the ek_ value; the browser POSTs its SDP offer to https://api.openai.com/v1/realtime/calls with that key and opens the oai-events data channel.
  4. Server or telephony bridge: open the WebSocket below with the API key and send session.update with session.type = realtime.
  5. Phone: point your SIP trunk at sip:<project_id>@sip.api.openai.com;transport=tls, add a realtime.call.incoming webhook, then POST /v1/realtime/calls/{call_id}/accept.
  6. Log response.done usage on every turn so you can see context growth and cache hit rates.

Endpoint

wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1 (WebSocket); https://api.openai.com/v1/realtime/calls (WebRTC SDP); sip:$PROJECT_ID@sip.api.openai.com;transport=tls (SIP)

Authentication

Authorization: Bearer <API key> for server WebSocket and SIP control. For browsers and mobile, mint an ephemeral key (ek_...) server side via POST /v1/realtime/client_secrets (expires_after.seconds 10 to 7200, default 600) and lock instructions/tools in that session config. Optional OpenAI-Safety-Identifier header on the server request is bound to the token.

Quick start javascript

import WebSocket from "ws"; // server side only: never ship the API key to a browser
const ws = new WebSocket("wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1", {
  headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },
});
ws.on("open", () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      type: "realtime",
      instructions: "You are a friendly support agent. Keep answers short.",
      audio: {
        input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: { type: "semantic_vad" } },
        output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" },
      },
    },
  }));
});
// Call for every ~100 ms chunk of 24 kHz mono PCM16 from your mic or phone bridge
export function sendAudio(pcm16) {
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "response.output_audio.delta") playPcm16(Buffer.from(ev.delta, "base64"));
  if (ev.type === "input_audio_buffer.speech_started") stopPlayback(); // barge-in: flush local audio
  if (ev.type === "response.done") console.log(ev.response.usage); // log tokens per turn
  if (ev.type === "error") console.error(ev.error);
});

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Context growth multiplies cost

Each response re-sends the whole conversation, including earlier audio (model audio is re-billed as audio input at $32/1M). A 10 minute call can cost 8x the naive per-minute figure with no cache hits. Changing instructions or tools mid-session, or truncating every turn, breaks the cache. Use token_limits.post_instructions with retention_ratio around 0.8, or delete or summarise old items.

Beta interface is gone

The OpenAI-Beta: realtime=v1 interface was removed May 12, 2026. GA requires session.type, nests audio config under session.audio.input/output, and renames events (response.output_audio.delta, response.output_text.delta, response.output_audio_transcript.delta). Old tutorials and Azure preview samples using response.audio.delta will silently get no audio.

Deprecation calendar

gpt-realtime, gpt-realtime-mini, gpt-4o-realtime and gpt-4o-mini-realtime shut down January 20, 2027. gpt-4o realtime previews already died May 7, 2026. Pin explicit model ids and test 2.1 before January.

Voice is locked after first audio

Once the model has emitted audio in a session the voice cannot be changed. Set it in the client secret or the first session.update.

60 minute hard cap without warning

Sessions end at 60 minutes and no warning event is sent. Track session age yourself and reconnect at a turn boundary, replaying a summary into the new session.

Reasoning effort trades latency and cost

gpt-realtime-2.x are reasoning models; higher reasoning effort increases both time to first audio and billed text output tokens ($24/1M). Keep effort low for chit-chat and raise it only for tool-heavy turns.

Mini models are less reliable with tools

OpenAI's own cost guide warns the mini model may follow instructions and call functions less reliably. Prototype on the full model, then measure regressions before switching.

Translation endpoint has its own retention rules

/v1/realtime/translations is not listed in the ZDR column of OpenAI's data controls table, unlike /v1/realtime. Check with OpenAI before sending regulated audio.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Max session length 60 minutes, no warning event before cutoff (third-party SDKs reconnect at ~50 min)
  • Context window 128,000 tokens; max output 32,000 (2.1 and 2.1-mini)
  • Default rate limits gpt-realtime-2.1: Build 400 RPM / 200,000 TPM, Launch 10,000 RPM / 4,000,000 TPM, Grow 20,000 RPM / 15,000,000 TPM
  • gpt-realtime-translate: Build 200, Launch 650, Grow 850 minutes of audio per minute
  • Ephemeral client secret lifetime 10 s to 7,200 s, default 600 s
  • Audio append chunk max 15 MB

Models and products

NameStatusNotes
gpt-realtime-2.1GACurrent flagship (released July 2026). Reasoning model with configurable reasoning effort, better alphanumerics, noise and interruption handling than gpt-realtime-2. 128k context, 32k max output. Text, audio and image input.
gpt-realtime-2.1-miniGADistilled, cheaper reasoning model. Same 128k context. Replacement for gpt-realtime-mini.
gpt-realtime-2GAMay 2026 release, same price as 2.1. Older but not yet on the deprecation list.
gpt-realtime-1.5GAOlder generation; text output $16/1M instead of $24.
gpt-realtimeDeprecatedShutdown January 20, 2027; replacement gpt-realtime-2.1.
gpt-realtime-miniDeprecatedShutdown January 20, 2027; replacement gpt-realtime-2.1-mini. The 2025-10-06 snapshot was already shut down July 23, 2026.
gpt-realtime-translateGASpeech-to-speech translation on /v1/realtime/translations, $0.034/min. Third-party reports say 70+ input and 13 output languages (not confirmed on the model page).
gpt-live-transcribe / gpt-realtime-whisper / gpt-transcribeGARealtime transcription session models (session.type = transcription). gpt-live-transcribe is the recommended one.
gpt-4o-realtime-preview (all snapshots), gpt-4o-mini-realtime-previewDeprecatedShut down May 7, 2026.

Docs and sources

Docs

Sources used

Not fully verified

Exact reasoning-effort levels and how to set them for gpt-realtime-2.x; latency numbers (only the third-party-reported 25 percent p95 improvement); the language lists for gpt-realtime-translate; codecs accepted on inbound Realtime SIP; whether Realtime sessions have a per-tier concurrent session cap (model page lists RPM/TPM only).

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?