Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| StepAudio Realtime StepFun | Phonic Phonic | OpenAI Realtime API OpenAI | |
|---|---|---|---|
| Category | Voice-to-voice | Voice-to-voice | Voice-to-voice |
| Status | GA | GA | GA |
| Est. per minute | n/a | $0.15 | $0.096 - 0.76 |
| How that was worked out | Audio tokens per second are not documented; measure usage on a test call. | Published starting price; volume pricing via sales. | Low = 1 min of user audio in (600 tokens x $32/1M) + 1 min of model audio out (1,200 tokens x $64/1M) on gpt-realtime-2.1, single turn, no caching, no text. High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. With perfect cache hits on history the same call is about $0.058/min. gpt-realtime-2.1-mini: low $0.030, high $0.237 (about $0.022 cached). |
| Pricing model | per-token | per-minute | per-token |
| Free tier | Not confirmed. | Not stated. | None. Free tier is not supported for realtime models; usage tiers are now named Build, Launch, Grow. |
| Connects by | WebSocket | WebSocket, Webhooks, SIP (Twilio, Telnyx), LiveKit plugin, Amazon Connect via Chime SIP | WebRTC, WebSocket, SIP |
| Audio in | pcm16 (sample rate not stated on the model page) | pcm_44100 (default), pcm_24000, pcm_16000, pcm_8000, mulaw_8000 | audio/pcm 24 kHz mono 16-bit LE (default), audio/pcmu and audio/pcma (G.711, 8 kHz) for telephony; WebRTC negotiates its own codec. Base64 chunks via input_audio_buffer.append, max 15 MB per chunk. |
| Audio out | pcm16 | Same options | audio/pcm 24 kHz mono 16-bit (default) or G.711 u-law/A-law; settable per session or per response. |
| Languages | Not listed; docs and examples are Chinese. | 51 languages (per docs index). | Multilingual; OpenAI does not publish a fixed list for gpt-realtime-2.1. Test your target languages and accents. |
| Latency (vendor claim) | Not published. | Vendor claim: sub-500 ms speech-in to speech-out. | Vendor claim (reported by third-party coverage of the July 2026 release): gpt-realtime-2.1 cut p95 latency by at least 25 percent versus earlier realtime models via better caching. No absolute number published. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Not stated. | HIPAA and SOC 2 compliance, 99.9% uptime SLA (vendor claim). | /v1/realtime is Zero Data Retention eligible; default abuse-monitoring logs kept 30 days, no application state stored. US and EU data residency for current realtime models (EU needs approved controls). SOC 2 and BAA availability are account-level OpenAI programs; confirm realtime coverage in your agreement. |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Flat $/min | - | $0.15 | - |
| Audio in $/1M tok | $1.5 | - | $32 |
| Audio out $/1M tok | $10 | - | $64 |
| Free tier | - | - | No |
| Free credit $ | - | - | - |
| Native S2S | Yes | Yes | Yes |
| Tools | - | Yes | Yes |
| Image in | - | - | Yes |
| Own LLM | No | No | No |
| Voices | - | - | 10 |
| Cloning | Yes | - | No |
| Latency ms | - | 500 | - |
| Languages | - | 51 | - |
| Context tokens | - | - | 128,000 |
| Max session min | - | - | 60 |
| Concurrency | - | - | - |
| WebRTC | No | No | Yes |
| WebSocket | Yes | Yes | Yes |
| Phone / SIP | No | Yes | Yes |
| HIPAA | - | Yes | Yes |
| SOC 2 | - | Yes | Yes |
| EU data | - | - | Yes |
| Open weights | No | No | No |
| High warnings | 1 | 0 | 3 |