Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

StepAudio Realtime
StepFun
Phonic
Phonic
OpenAI Realtime API
OpenAI
CategoryVoice-to-voiceVoice-to-voiceVoice-to-voice
StatusGAGAGA
Est. per minuten/a$0.15$0.096 - 0.76
How that was worked outAudio tokens per second are not documented; measure usage on a test call.Published starting price; volume pricing via sales.Low = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached).
Pricing modelper-tokenper-minuteper-token
Free tierNot confirmed.Not stated.None. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow.
Connects byWebSocketWebSocket, Webhooks, SIP (Twilio, Telnyx), LiveKit plugin, Amazon Connect via Chime SIPWebRTC, WebSocket, SIP
Audio inpcm16 (sample rate not stated on the model page)pcm_44100 (default), pcm_24000, pcm_16000, pcm_8000, mulaw_8000audio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk.
Audio outpcm16Same optionsaudio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response.
LanguagesNot listed; docs and examples are Chinese.51 languages (per docs index).Multilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents.
Latency (vendor claim)Not published.Vendor claim: sub-500 ms speech-in to speech-out.Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published.
Key limits
  • Account rate tiers on the open platform range from V0 (under $15 top-up: 5 concurrency, 100 RPM) to V4 ($1,500+: 130 concurrency, 2,600 RPM); not stated whether these apply to realtime
  • Session and context limits not documented on the model page
  • Rate limits per organization: 500 requests/second shared across /sts/ws and SIP outbound; 5 requests/second for /conversations/outbound_call
  • Max session length 60 minutes, no warning event before cutoff (third-party SDKs reconnect at ~50 min)
  • Context window 128,000 tokens; max output 32,000 (2.1 and 2.1-mini)
  • Default rate limits gpt-realtime-2.1: Build 400 RPM / 200,000 TPM, Launch 10,000 RPM / 4,000,000 TPM, Grow 20,000 RPM / 15,000,000 TPM
  • gpt-realtime-translate: Build 200, Launch 650, Grow 850 minutes of audio per minute
High-severity warnings
  • Cost per minute is unknowable from docs
  • None
  • Context growth multiplies cost
  • Beta interface is gone
  • Deprecation calendar
ComplianceNot stated.HIPAA and SOC 2 compliance, 99.9% uptime SLA (vendor claim)./v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement.
Self-hostableNoNoNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
Flat $/min-$0.15-
Audio in $/1M tok$1.5-$32
Audio out $/1M tok$10-$64
Free tier--No
Free credit $---
Native S2SYesYesYes
Tools-YesYes
Image in--Yes
Own LLMNoNoNo
Voices--10
CloningYes-No
Latency ms-500-
Languages-51-
Context tokens--128,000
Max session min--60
Concurrency---
WebRTCNoNoYes
WebSocketYesYesYes
Phone / SIPNoYesYes
HIPAA-YesYes
SOC 2-YesYes
EU data--Yes
Open weightsNoNoNo
High warnings103