Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

LFM2.5-Audio-1.5B
Liquid AI
Moshi
Kyutai
Unmute
Kyutai
CategoryOpen modelsOpen modelsOpen models
StatusGAGAGA
Est. per minuten/an/an/a
How that was worked outNo licence fee; you pay for GPU time. No licence fee; you pay for GPU time. No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint.
Pricing modelfreefreefree
Free tierOpen weightsOpen weightsOpen weights
Connects byPython package (liquid-audio), llama.cpp GGUF, ONNXWebSocket (built-in web server and client)WebSocket (OpenAI-Realtime-like backend), Web frontend
Audio inSpeech24 kHz via the Mimi codec (12.5 Hz frames, 80 ms)Browser mic via the bundled frontend
Audio outSpeech (interleaved with text)24 kHz Mimi-decoded speechStreamed TTS audio
LanguagesEnglish (Japanese variant available).English (not stated on the repo page; the released models were trained for English).English and French for STT/TTS (per model names); LLM language depends on your model.
Latency (vendor claim)Designed for low-latency real-time conversation; no figure.Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU.Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs.
Key limits
  • Small model: limited knowledge/reasoning
  • English only (main model)
  • Small 7B text backbone: limited knowledge and reasoning
  • No built-in tool/function calling
  • Context length not stated on the repo
  • No built-in tool calling (wrap your LLM server to add it)
  • Linux x86_64 only (Windows via WSL); no aarch64 or native Mac
High-severity warnings
  • Revenue cap in the licence
  • Not an agent brain
  • None
ComplianceSelf-hosted.Self-hosted; compliance is your responsibility.Self-hosted; your responsibility.
Self-hostableYesYesYes
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
TypeVoice-to-voiceVoice-to-voiceVoice-to-voice
Params B1.57.7-
VRAM GB-2416
CPU okYesNoNo
LicenceCustomCC-BY-4.0CC-BY-4.0
CommercialNoYesYes
Full duplexNoYesNo
Streaming-YesYes
Latency ms-200750
Languages112
High warnings110