Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Chatterbox (Turbo, Nano, Multilingual V3) Resemble AI | Moshi Kyutai | Unmute Kyutai | |
|---|---|---|---|
| Category | Open models | Open models | Open models |
| Status | GA | GA | GA |
| Est. per minute | n/a | n/a | n/a |
| How that was worked out | Self-hosted: cost is your GPU/CPU time, not per character | No licence fee; you pay for GPU time. | No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint. |
| Pricing model | free | free | free |
| Free tier | Open weights | Open weights | Open weights |
| Connects by | HTTP chunked | WebSocket (built-in web server and client) | WebSocket (OpenAI-Realtime-like backend), Web frontend |
| Audio in | Text plus optional reference audio prompt | 24 kHz via the Mimi codec (12.5 Hz frames, 80 ms) | Browser mic via the bundled frontend |
| Audio out | Waveform at model sample rate (model.sr) | 24 kHz Mimi-decoded speech | Streamed TTS audio |
| Languages | Turbo/Nano: English; Multilingual V3: 23+ | English (not stated on the repo page; the released models were trained for English). | English and French for STT/TTS (per model names); LLM language depends on your model. |
| Latency (vendor claim) | Hosted Resemble service claims sub-200 ms; self-hosted depends on GPU. No official streaming API in the library. | Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU. | Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Your own deployment; no vendor data processing | Self-hosted; compliance is your responsibility. | Self-hosted; your responsibility. |
| Self-hostable | Yes | Yes | Yes |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Type | Text-to-speech | Voice-to-voice | Voice-to-voice |
| Params B | 0.35 | 7.7 | - |
| VRAM GB | - | 24 | 16 |
| CPU ok | Yes | No | No |
| Licence | MIT | CC-BY-4.0 | CC-BY-4.0 |
| Commercial | Yes | Yes | Yes |
| Full duplex | - | Yes | No |
| Streaming | No | Yes | Yes |
| Latency ms | 103 | 200 | 750 |
| Languages | 23 | 1 | 2 |
| High warnings | 1 | 1 | 0 |