Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

MiniCPM-o 4.5
OpenBMB (ModelBest / Tsinghua)
Moshi
Kyutai
Unmute
Kyutai
CategoryOpen modelsOpen modelsOpen models
StatusGAGAGA
Est. per minuten/an/an/a
How that was worked outNo licence fee; you pay for GPU time. No licence fee; you pay for GPU time. No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint.
Pricing modelfreefreefree
Free tierOpen weightsOpen weightsOpen weights
Connects byPython (Transformers), vLLM, SGLang, llama.cpp-omni, OllamaWebSocket (built-in web server and client)WebSocket (OpenAI-Realtime-like backend), Web frontend
Audio inStreaming speech (and video frames)24 kHz via the Mimi codec (12.5 Hz frames, 80 ms)Browser mic via the bundled frontend
Audio outStreaming speech24 kHz Mimi-decoded speechStreamed TTS audio
LanguagesReal-time speech conversation in English and Chinese; text in 30+ languages.English (not stated on the repo page; the released models were trained for English).English and French for STT/TTS (per model names); LLM language depends on your model.
Latency (vendor claim)Vendor efficiency table: time to first token 0.6 s; decoding 154 tok/s (bf16) and 212 tok/s (int4).Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU.Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs.
Key limits
  • Half-duplex speech streaming marked under development
  • Speech conversation is bilingual (EN/ZH) only
  • Small 7B text backbone: limited knowledge and reasoning
  • No built-in tool/function calling
  • Context length not stated on the repo
  • No built-in tool calling (wrap your LLM server to add it)
  • Linux x86_64 only (Windows via WSL); no aarch64 or native Mac
High-severity warnings
  • None
  • Not an agent brain
  • None
ComplianceSelf-hosted.Self-hosted; compliance is your responsibility.Self-hosted; your responsibility.
Self-hostableYesYesYes
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
TypeVoice-to-voiceVoice-to-voiceVoice-to-voice
Params B97.7-
VRAM GB112416
CPU ok-NoNo
LicenceApache-2.0CC-BY-4.0CC-BY-4.0
CommercialYesYesYes
Full duplexYesYesNo
StreamingYesYesYes
Latency ms600200750
Languages212
High warnings010