Azure OpenAI GPT Realtime API (Microsoft Foundry)
The OpenAI gpt-realtime models hosted in your Azure subscription, with Entra ID auth, Azure networking and Data Zone options. Suits teams that must stay inside Azure contracts or regions.
Overview
Best for: Enterprises already on Azure that need private networking, Entra ID, Microsoft contracts and US/EU Data Zone processing for OpenAI realtime voice.
At a glance
Global prices for gpt-realtime-2.1; Data Zone is 1.1x ($35.20/$70.40). Microsoft quotes transport latency only (WebRTC ~100 ms, WebSocket ~200 ms). Docs still state a 32k input token limit that conflicts with the 128k model context. Generic Azure free-account credits may apply.
PCM16 mono 24 kHz recommended (send ~100 ms chunks); G.711 supported per the shared OpenAI event model; WebRTC negotiates codecs.
PCM16 24 kHz (same options as OpenAI).
Same models as OpenAI; Microsoft advises validating languages with production-like audio and passing ISO-639-1 hints for transcription.
Same 10 OpenAI voices (marin and cedar recommended). For Azure neural or custom voices use Voice Live instead.
Microsoft guidance (transport only, not model time): WebRTC ~100 ms, WebSocket ~200 ms.
Global deployments served from East US 2 and Sweden Central for WebRTC/realtime; Data Zone (US, EU) keeps processing inside the zone. Check the region availability page for each model.
Covered by Azure OpenAI enterprise terms (Microsoft Products and Services DPA; HIPAA BAA via Microsoft for in-scope Azure services). Data Zone deployments keep processing within the US or EU zone. Content filtering applies.
Features
- function calling and remote MCP servers
- server_vad, semantic_vad or manual turn handling
- image input
- out-of-band responses
- Entra ID keyless auth and managed identity
- ephemeral tokens via /openai/v1/realtime/client_secrets
- SIP telephony
- Global, Data Zone (US/EU) deployment types
Pricing
| What | Price | Unit |
|---|---|---|
| gpt-realtime-2.1 Global audio input / output | $32.00 / $64.00 | per 1M tokens |
| gpt-realtime-2.1 Global text input / output | $4.00 / $24.00 | per 1M tokens |
| gpt-realtime-2.1 Global image input | $5.00 | per 1M tokens |
| gpt-realtime-2.1 Data Zone audio input / output | $35.20 / $70.40 | per 1M tokens |
| gpt-realtime-2.1-mini Global audio input / output | $10.00 / $20.00 | per 1M tokens |
| gpt-realtime-2.1-mini Data Zone audio input / output | $11.00 / $22.00 | per 1M tokens |
| gpt-realtime-2 Global | $32.00 / $64.00 audio, $4.00 / $24.00 text | per 1M tokens |
| gpt-realtime-1.5 Global / Data Zone audio | $32.00 / $64.00 and $35.20 / $70.40 | per 1M tokens |
| gpt-realtime-translate (Global) | $2.04 | per hour |
| gpt-realtime-whisper (Global) | $1.02 | per hour |
Low = gpt-realtime-2.1 Global, 1 min user audio (600 tokens) + 1 min model audio (1,200 tokens), single turn, no cache. Data Zone single turn $0.106. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. High shown at Data Zone rates ($0.76 Global); with full cache hits about $0.06/min.
Same as OpenAI: about 10 tokens/s input audio, 20 tokens/s output audio (Microsoft Voice Live token table for Azure OpenAI models).
Free tier: No free tier for realtime models (Azure free account credits can apply).
Source: prices.azure.com
Setup
- Create an Azure subscription and a Microsoft Foundry resource in a supported region (East US 2 or Sweden Central are the documented realtime regions).
- In the Foundry portal deploy a realtime model (e.g. gpt-realtime-2.1) as Global or Data Zone; note the deployment name.
- Prefer Entra ID: assign Cognitive Services OpenAI User and request tokens for scope https://ai.azure.com/.default; otherwise copy the resource key.
- Server or telephony: open the GA WebSocket (/openai/v1/realtime) with the deployment name as model. Do not add api-version.
- Browser: your token service calls https://<resource>.openai.azure.com/openai/v1/realtime/client_secrets and the browser posts its SDP to https://<resource>.openai.azure.com/openai/v1/realtime/calls with the ephemeral token.
- Request quota increases early: 100k TPM is small for realtime because of context re-billing.
Endpoint
wss://<resource>.openai.azure.com/openai/v1/realtime?model=<deployment> (WebSocket); https://<resource>.openai.azure.com/openai/v1/realtime/calls (WebRTC); https://<resource>.openai.azure.com/openai/v1/realtime/client_secrets (ephemeral tokens)
Authentication
api-key header or Authorization: Bearer <Entra ID token> on the server. Browsers get an ephemeral token from your token service; the more secure pattern proxies the SDP exchange so the browser never holds even the ephemeral token.
Quick start javascript
import WebSocket from "ws"; // server side; browsers use /openai/v1/realtime/client_secrets + WebRTC
const RESOURCE = process.env.AZURE_OPENAI_RESOURCE; // e.g. my-foundry-resource
const DEPLOYMENT = process.env.AZURE_OPENAI_DEPLOYMENT; // your deployment name, not the model name
const ws = new WebSocket(
`wss://${RESOURCE}.openai.azure.com/openai/v1/realtime?model=${DEPLOYMENT}`,
{ headers: { "api-key": process.env.AZURE_OPENAI_API_KEY } } // or Authorization: Bearer <Entra token>
);
ws.on("open", () => {
ws.send(JSON.stringify({
type: "session.update",
session: {
type: "realtime",
instructions: "You are a friendly support agent. Keep answers short.",
audio: {
input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: { type: "server_vad" } },
output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" },
},
},
}));
});
export function sendAudio(pcm16) { // 24 kHz mono PCM16, ~100 ms chunks
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.output_audio.delta") playPcm16(Buffer.from(ev.delta, "base64"));
if (ev.type === "input_audio_buffer.speech_started") stopPlayback();
if (ev.type === "error") console.error(ev.error);
});
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Preview endpoints and samples are deprecated
Old samples use /openai/realtimeapi/sessions, https://<region>.realtimeapi-preview.ai.azure.com/v1/realtimertc, api-version query strings and beta event names (response.audio.delta). The GA path is /openai/v1/realtime/... with no api-version; mixing them gives 404s or silent audio.
Low default TPM
The Tier 1 table shows gpt-realtime at 100,000 TPM. Because every turn re-sends history, a handful of concurrent long calls can hit 429s. Plan quota requests and back-off before launch.
Data Zone costs 10 percent more
Data Zone meters are exactly 1.1x Global (e.g. $35.20 vs $32.00 audio input). Use Data Zone only when residency requires it.
Model is a deployment, not a name
The model query parameter is your deployment name. Upgrading to a new model version means a new or updated deployment, and retired versions follow Azure's own retirement schedule, not OpenAI's.
Entra auth gotcha
Keyless auth fails if the AZURE_OPENAI_API_KEY environment variable is set; the SDK picks the key. Unset it when using DefaultAzureCredential.
Region-limited availability
Realtime Global deployments are documented for East US 2 and Sweden Central only. Latency from other continents can be noticeably worse than OpenAI's global edge.
Inconsistent limits in docs
The how-to page states both a 60 minute session cap and a 32,000 input / 4,096 output token limit that predates 128k-context models. Test the real limits on your deployment.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Max session duration 60 minutes (monitor expires_at in session.created)
- Default gpt-realtime quota in the Tier 1 table: 200 RPM and 100,000 TPM (GlobalStandard); higher tiers raise it (e.g. 300 RPM / 150,000 TPM)
- GPT-Live on Azure: concurrent sessions per subscription Default 10, Tier 1 25, Tier 2 50, Tier 3 200, Tier 4 300, Tier 5 500
- Realtime quota is separate from chat completions quota
- Doc still states 32,000 input / 4,096 output tokens for the Realtime API, which conflicts with the 128k model context on OpenAI
Models and products
| Name | Status |
|---|---|
| gpt-realtime-2.1 (2026-07-07) | GA |
| gpt-realtime-2.1-mini (2026-07-07) | GA |
| gpt-realtime-2 (2026-05-07) | GA |
| gpt-realtime-1.5 (2026-02-23) | GA |
| gpt-realtime (2025-08-28), gpt-realtime-mini (2025-10-06, 2025-12-15) | GA |
| gpt-4o-realtime-preview / gpt-4o-mini-realtime-preview (2024-12-17) | Preview |
| gpt-realtime-translate, gpt-realtime-whisper (2026-05-06), gpt-live-transcribe (2026-07-29) | GA |
| gpt-live-1 (GPT-Live) | Preview |
Docs and sources
Docs
Sources used
- learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio
- learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio-webrtc
- learn.microsoft.com/en-us/azure/foundry/openai/how-to/realtime-audio-websockets
- learn.microsoft.com/en-us/azure/foundry/openai/quotas-limits
- prices.azure.com/api/retail/prices
- learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live
Full raw WebSocket URL (docs show only the SDK base /openai/v1 with model=<deployment>; the ?model= form is inferred); Azure SIP URI format; per-model region list beyond East US 2 / Sweden Central; whether gpt-live-1 is generally available on Azure; retail prices are from the eastus2 Retail Prices API because the public pricing page renders prices client-side.