Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Cartesia Managed Agents (Line) Cartesia | OpenAI GPT-Live API OpenAI | Ultravox Realtime Ultravox (formerly Fixie.ai) | |
|---|---|---|---|
| Category | Voice-to-voice | Voice-to-voice | Voice-to-voice |
| Status | GA | GA | GA |
| Est. per minute | $0.06 - 0.09 | $0.05 | $0.05 - 0.055 |
| How that was worked out | $0.06 base plus $0.014 telephony on Cartesia numbers plus LLM tokens (small models add little; own estimate). | Low = voice layer only, $0.05 per minute of session duration with no delegated work. High depends entirely on the backend model, how often the model delegates and tool fees; budget per-minute session cost plus backend token spend measured in your own tests. | $0.05 call rate plus $0.005 SIP when using SIP; LLM and TTS included. |
| Pricing model | per-minute | per-minute | per-minute |
| Free tier | Subscription plans include monthly credits; free LLM promo ended 2026-10-01. | Free tier not supported. | 30 free call minutes; playground calls free. |
| Connects by | WebSocket, Phone numbers (Cartesia, Twilio import), SIP trunking | WebRTC, WebSocket, SIP | WebRTC, WebSocket, SIP, Twilio, Telnyx, Plivo, Exotel |
| Audio in | pcm_16000, pcm_24000, pcm_44100 (16-bit) or mulaw_8000, headerless mono | WebSocket: audio/pcm 24 kHz (default) or 16 kHz mono 16-bit LE, audio/pcmu or audio/pcma 8 kHz; raw bytes base64, even byte length, no container. Format fixed at session start. WebRTC and SIP negotiate codecs (outbound SIP needs Opus + SDES-SRTP). | Raw PCM s16le for server WebSocket (sample rate set in the call medium) |
| Audio out | Streamed audio_output events (speaking-pace delivery option) | Same format as input (one setting covers both). | PCM for server WebSocket; WebRTC handles codecs automatically |
| Languages | Configurable per agent; list not read. | Not listed on the model page. | Multilingual; no list on the pages read. |
| Latency (vendor claim) | Not stated on the pages read. | Vendor claim: improves Full Duplex Bench score by 30 percentage points over gpt-realtime-2.1 (reported via third-party coverage). No millisecond figure published. | Vendor says audio-native processing is faster and robust to transcription errors; no figure on the pages read. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Not stated on the pages read. | /v1/live/sessions is ZDR eligible with limitations (store forced false, no forking or recording download). Abuse-monitoring logs 30 days. US and EU data residency. | Not stated on the pricing or FAQ pages (a trust portal is linked). |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Flat $/min | $0.06 | $0.05 | $0.05 |
| Audio in $/1M tok | - | - | - |
| Audio out $/1M tok | - | - | - |
| Free tier | - | No | Yes |
| Free credit $ | - | - | - |
| Native S2S | No | Yes | No |
| Tools | Yes | Yes | Yes |
| Image in | - | No | - |
| Own LLM | - | Yes | - |
| Voices | - | 12 | - |
| Cloning | - | - | Yes |
| Latency ms | - | - | - |
| Languages | - | - | - |
| Context tokens | - | 128,000 | - |
| Max session min | - | - | - |
| Concurrency | - | 50 | 5 |
| WebRTC | No | Yes | Yes |
| WebSocket | Yes | Yes | Yes |
| Phone / SIP | Yes | Yes | Yes |
| HIPAA | - | - | - |
| SOC 2 | - | - | - |
| EU data | - | Yes | - |
| Open weights | No | No | Yes |
| High warnings | 1 | 2 | 1 |