Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Hume EVI (Empathic Voice Interface) Hume AI | Deepgram Voice Agent API Deepgram | OpenAI GPT-Live API OpenAI | |
|---|---|---|---|
| Category | Voice-to-voice | Voice-to-voice | Voice-to-voice |
| Status | Deprecated | GA | GA |
| Est. per minute | $0.04 - 0.07 | $0.041 - 0.163 | $0.05 |
| How that was worked out | Third-party reported EVI 3 overage rates; EVI 4-mini reportedly about half. Supplemental LLM cost may be extra. | Official tier rates; BYO tiers add your own LLM/TTS bills on top. | Low = voice layer only, $0.05 per minute of session duration with no delegated work. High depends entirely on the backend model, how often the model delegates and tool fees; budget per-minute session cost plus backend token spend measured in your own tests. |
| Pricing model | subscription | per-minute | per-minute |
| Free tier | 5 EVI minutes per month (third-party data). | $200 one-time credit for new accounts. | Free tier not supported. |
| Connects by | WebSocket | WebSocket | WebRTC, WebSocket, SIP |
| Audio in | WebM (browser) or linear16 PCM (e.g. 44.1 kHz mono, declared in session_settings); no mu-law | linear16 default at 16 kHz (other encodings/rates configurable) | WebSocket: audio/pcm 24 kHz (default) or 16 kHz mono 16-bit LE, audio/pcmu or audio/pcma 8 kHz; raw bytes base64, even byte length, no container. Format fixed at session start. WebRTC and SIP negotiate codecs (outbound SIP needs Opus + SDES-SRTP). |
| Audio out | base64 WAV in audio_output messages | linear16 (e.g. 24 kHz), optional container; other encodings configurable | Same format as input (one setting covers both). |
| Languages | EVI 3: English. EVI 4-mini: English, Japanese, Korean, Spanish, French, Portuguese, Italian, German, Russian, Hindi, Arabic. | English by default; multilingual via Nova 'multi' or flux-general-multi with language hints, and multilingual TTS providers. | Not listed on the model page. |
| Latency (vendor claim) | No figure on the overview page. | Not stated on the pages read; the API reports per-turn latency metrics (time to first token, TTS time to first byte). | Vendor claim: improves Full Duplex Bench score by 30 percentage points over gpt-realtime-2.1 (reported via third-party coverage). No millisecond figure published. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Not re-verified; irrelevant after shutdown. | Not re-verified this session (Deepgram markets SOC 2 and HIPAA readiness; check the trust page). | /v1/live/sessions is ZDR eligible with limitations (store forced false, no forking or recording download). Abuse-monitoring logs 30 days. US and EU data residency. |
| Self-hostable | No | No | No |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Flat $/min | $0.07 | $0.075 | $0.05 |
| Audio in $/1M tok | - | - | - |
| Audio out $/1M tok | - | - | - |
| Free tier | Yes | Yes | No |
| Free credit $ | - | $200 | - |
| Native S2S | Yes | No | Yes |
| Tools | - | Yes | Yes |
| Image in | - | - | No |
| Own LLM | Yes | Yes | Yes |
| Voices | - | - | 12 |
| Cloning | Yes | - | - |
| Latency ms | - | - | - |
| Languages | 11 | - | - |
| Context tokens | - | - | 128,000 |
| Max session min | 30 | - | - |
| Concurrency | - | 45 | 50 |
| WebRTC | No | No | Yes |
| WebSocket | Yes | Yes | Yes |
| Phone / SIP | No | No | Yes |
| HIPAA | - | - | - |
| SOC 2 | - | - | - |
| EU data | - | - | Yes |
| Open weights | No | No | No |
| High warnings | 2 | 2 | 2 |