Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Kokoro-82M
hexgrad (community)
Moshi
Kyutai
Unmute
Kyutai
CategoryOpen modelsOpen modelsOpen models
StatusGAGAGA
Est. per minuten/an/an/a
How that was worked outSelf-hosted: cost is your GPU/CPU time, not per characterNo licence fee; you pay for GPU time. No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint.
Pricing modelfreefreefree
Free tierOpen weightsOpen weightsOpen weights
Connects byHTTP chunkedWebSocket (built-in web server and client)WebSocket (OpenAI-Realtime-like backend), Web frontend
Audio inText (phonemized via misaki/espeak)24 kHz via the Mimi codec (12.5 Hz frames, 80 ms)Browser mic via the bundled frontend
Audio out24 kHz float audio from the Python library; community servers expose wav/mp3/pcm24 kHz Mimi-decoded speechStreamed TTS audio
Languages8 (American/British English, plus others listed in VOICES.md)English (not stated on the repo page; the released models were trained for English).English and French for STT/TTS (per model names); LLM language depends on your model.
Latency (vendor claim)No official TTFB figure; generates per sentence segment.Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU.Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs.
Key limits
  • No native text-input streaming: you feed sentences
  • Quality drops on unusual words (espeak fallback)
  • Small 7B text backbone: limited knowledge and reasoning
  • No built-in tool/function calling
  • Context length not stated on the repo
  • No built-in tool calling (wrap your LLM server to add it)
  • Linux x86_64 only (Windows via WSL); no aarch64 or native Mac
High-severity warnings
  • None
  • Not an agent brain
  • None
ComplianceYour own deployment; no vendor data processingSelf-hosted; compliance is your responsibility.Self-hosted; your responsibility.
Self-hostableYesYesYes
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
TypeText-to-speechVoice-to-voiceVoice-to-voice
Params B0.0827.7-
VRAM GB-2416
CPU okYesNoNo
LicenceApache-2.0CC-BY-4.0CC-BY-4.0
CommercialYesYesYes
Full duplex-YesNo
StreamingNoYesYes
Latency ms-200750
Languages812
High warnings010