GA OpenAI

OpenAI Realtime transcription sessions

Transcription-only sessions on the Realtime API (WebSocket or WebRTC). gpt-live-transcribe and gpt-realtime-whisper stream deltas; gpt-transcribe and gpt-4o-transcribe transcribe each committed turn.

Est. per minute$0.003 - 0.017
2 high-severity warnings

Overview

Best for: Apps already on the OpenAI Realtime stack that want transcription in the same session model, or browser capture via WebRTC.

At a glance

$/hour$1.02
Free tierNo
Live speakersNo
KeytermsYes
PartialsYes
WebRTCYes
WebSocketYes
gRPCNo
Self-hostNo
Open weightsNo

Price is true delta streaming (gpt-live-transcribe or gpt-realtime-whisper, $0.017/min). gpt-transcribe is $0.27/hr but only returns text after a commit. gpt-live-transcribe has no server VAD; you commit turns yourself. Diarization model is file-only. Rate limits depend on account tier.

Audio in

audio/pcm at 24 kHz in the documented example (16-bit, base64 in input_audio_buffer.append); G.711 formats historically supported on Realtime (not re-verified).

Audio out

n/a

Languages

Multilingual (count not published on the pages read); gpt-live-transcribe accepts a 'languages' hint list.

Latency

Vendor: 'low-latency' with tunable delay setting; no number published.

Regions

OpenAI API global; data residency options not checked for these models.

Compliance

OpenAI API data controls apply (API data not used for training by default per OpenAI policy); BAA availability not re-verified for these models.

Features

  • transcript deltas
  • final transcript per committed turn
  • client-side commit (input_audio_buffer.commit)
  • prompt / keywords / language hints
  • delay setting (minimal to xhigh) on gpt-live-transcribe
  • WebRTC for browsers with ephemeral client secrets

Pricing

WhatPriceUnitNotes
gpt-live-transcribe$0.017per minute of audioRealtime audio duration pricing.
gpt-realtime-whisper$0.017per minute of audio
gpt-transcribe$0.0045per minutePer committed turn / file.
gpt-4o-transcribe$2.50 in / $10.00 outper 1M tokensOpenAI estimate ~$0.006/min.
gpt-4o-mini-transcribe$1.25 in / $5.00 outper 1M tokensOpenAI estimate ~$0.003/min.
Whisper (file API)$0.006per minuteBatch only, not streaming.
How the per-minute estimate was worked out

$0.003/min (4o-mini-transcribe estimate, per turn) to $0.017/min (true streaming deltas with gpt-live-transcribe or gpt-realtime-whisper).

Free tier: None specific to transcription.

Source: developers.openai.com

Setup

  1. Create an API key.
  2. Server: open a Realtime WebSocket with Authorization: Bearer <key>; the OpenAI cookbook pattern uses wss://api.openai.com/v1/realtime?intent=transcription.
  3. Send session.update with session.type='transcription', audio.input.format, transcription.model and turn_detection (null for gpt-live-transcribe).
  4. Append base64 PCM with input_audio_buffer.append, commit each turn with input_audio_buffer.commit (your own VAD).
  5. Browser: mint an ephemeral client secret on your server and connect via WebRTC.

Endpoint

wss://api.openai.com/v1/realtime?intent=transcription (route listed as v1/realtime/transcription_sessions on model pages)

Authentication

Authorization: Bearer <API_KEY> server side; ephemeral client secrets for browser WebRTC.

Quick start javascript

import WebSocket from "ws"; // npm i ws
import fs from "fs";

const ws = new WebSocket("wss://api.openai.com/v1/realtime?intent=transcription", {
  headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },
});
ws.on("open", async () => {
  ws.send(JSON.stringify({ type: "session.update", session: { type: "transcription",
    audio: { input: { format: { type: "audio/pcm", rate: 24000 },
      transcription: { model: "gpt-live-transcribe" }, turn_detection: null } } } }));
  const pcm = fs.readFileSync("audio_24k_mono_s16le.raw");
  for (let i = 0; i < pcm.length; i += 4800) { // 100 ms at 24 kHz
    ws.send(JSON.stringify({ type: "input_audio_buffer.append",
      audio: pcm.subarray(i, i + 4800).toString("base64") }));
    await new Promise((r) => setTimeout(r, 100));
  }
  ws.send(JSON.stringify({ type: "input_audio_buffer.commit" })); // end of turn
});
ws.on("message", (data) => {
  const ev = JSON.parse(data);
  if (ev.type === "conversation.item.input_audio_transcription.delta") process.stdout.write(ev.delta);
  if (ev.type === "conversation.item.input_audio_transcription.completed") console.log("\nFINAL:", ev.transcript);
  if (ev.type === "error") console.error(ev.error);
});

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No server VAD on gpt-live-transcribe

The recommended streaming model rejects server_vad and semantic_vad. You must run client-side VAD and send input_audio_buffer.commit at the end of each turn, or you never get a final transcript.

True streaming is ~3x the cost of competitors

Delta-streaming models (gpt-live-transcribe, gpt-realtime-whisper) are $0.017/min, roughly 3-8x Deepgram, AssemblyAI or Soniox list prices. gpt-transcribe at $0.0045/min only returns text after a commit.

Model names churned in 2026

gpt-realtime-whisper (May 2026) and gpt-live-transcribe / gpt-transcribe (July 2026) arrived within months; third-party guides disagree on which to use. Follow the realtime transcription guide and pin a model.

Out-of-order completions

Completed events from different turns can arrive out of order. Key transcripts by item_id rather than appending in arrival order.

Azure lags OpenAI

A May 2026 Microsoft Q&A thread reports Azure did not list gpt-realtime-whisper for live input transcription while it worked on OpenAI directly. Check model availability per Azure region.

Language param is plural

gpt-live-transcribe takes 'languages' (a list); sending the older singular 'language' field is not supported.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Rate limits depend on account usage tier (not captured).
  • Completion events from different turns can arrive out of order; match by item_id.

Models and products

NameStatusNotes
gpt-live-transcribeGA (recommended in realtime guide)Streams deltas, tunable delay (minimal..xhigh), prompt, keywords, multiple language hints. Does NOT support server_vad/semantic_vad: you must commit turns yourself. No language detection output.
gpt-realtime-whisperGAStreaming STT priced by audio duration; supported on v1/realtime/transcription_sessions.
gpt-transcribeGAHigh-accuracy model for committed turns (starts after commit) and files; returns detected languages.
gpt-4o-transcribeGA (older)Token-priced; usable as transcription model in realtime sessions.
gpt-4o-mini-transcribeGA (older)Cheapest token-priced option.
gpt-4o-transcribe-diarizeGA (file transcription)Speaker labels; file API, not a live streaming model.

Docs and sources

Docs

Sources used

Not fully verified

Exact WebSocket URL for transcription sessions (the ?intent=transcription query comes from a cookbook pattern, not the guide), rate limits, supported language count, and whether gpt-realtime-whisper is being superseded by gpt-live-transcribe.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?