Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Qwen3-Omni-30B-A3B (open weights)
Alibaba Qwen
Moshi
Kyutai
Unmute
Kyutai
CategoryOpen modelsOpen modelsOpen models
StatusGAGAGA
Est. per minuten/an/an/a
How that was worked outNo licence fee; you pay for GPU time. Hosted equivalents are on Alibaba Model Studio (see the Qwen-Omni Realtime entry).No licence fee; you pay for GPU time. No licence fee; you pay for GPU time. Plus LLM cost if you use a paid LLM endpoint.
Pricing modelfreefreefree
Free tierOpen weightsOpen weightsOpen weights
Connects byPython (Transformers), vLLM (text output only today)WebSocket (built-in web server and client)WebSocket (OpenAI-Realtime-like backend), Web frontend
Audio inSpeech in 19 languages (18 listed)24 kHz via the Mimi codec (12.5 Hz frames, 80 ms)Browser mic via the bundled frontend
Audio outSpeech in 10 languages24 kHz Mimi-decoded speechStreamed TTS audio
LanguagesText 119 languages; speech input about 19; speech output 10 (EN, ZH, FR, DE, RU, IT, ES, PT, JA, KO).English (not stated on the repo page; the released models were trained for English).English and French for STT/TTS (per model names); LLM language depends on your model.
Latency (vendor claim)Not given as a number on the model card.Vendor claim: 160 ms theoretical, as low as 200 ms in practice on an L4 GPU.Vendor: TTS latency about 750 ms on a single L40S, about 450 ms with services on separate GPUs.
Key limits
  • vLLM serving supports only the thinker (no audio output) at the time of the card
  • No turn-key realtime/duplex server in the release
  • Small 7B text backbone: limited knowledge and reasoning
  • No built-in tool/function calling
  • Context length not stated on the repo
  • No built-in tool calling (wrap your LLM server to add it)
  • Linux x86_64 only (Windows via WSL); no aarch64 or native Mac
High-severity warnings
  • No speech output from vLLM
  • Big GPU bill
  • Not an agent brain
  • None
ComplianceSelf-hosted.Self-hosted; compliance is your responsibility.Self-hosted; your responsibility.
Self-hostableYesYesYes
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
TypeVoice-to-voiceVoice-to-voiceVoice-to-voice
Params B357.7-
VRAM GB792416
CPU okNoNoNo
LicenceApache-2.0CC-BY-4.0CC-BY-4.0
CommercialYesYesYes
Full duplexNoYesNo
Streaming-YesYes
Latency ms-200750
Languages1012
High warnings210