Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

whisper_streaming and SimulStreaming (UFAL)
Charles University UFAL (open source)
NVIDIA Nemotron ASR Streaming / Parakeet (Riva, NIM, NeMo)
NVIDIA
Kyutai STT (Delayed Streams Modeling)
Kyutai
CategoryOpen modelsOpen modelsOpen models
StatusGAGAGA
Est. per minute$0$0$0
How that was worked outSelf-host compute only.Self-host compute only; cost depends on GPU price and streams per GPU.Self-host compute only.
Pricing modelfreefreefree
Free tierOpen source.Open weights; hosted API trial on build.nvidia.com.Open weights.
Connects byTCP/stdin server scripts (research code)gRPC (Riva/NIM), WebSocket (Riva realtime client), Python (NeMo)WebSocket (moshi-server), Python (PyTorch, MLX)
Audio in16 kHz mono audio via the provided server/clients.16 kHz mono PCM typical for Riva streaming (not re-verified per model).Handled by the provided scripts (mic or file); server streams PCM over WebSocket.
Audio outn/an/an/a
LanguagesWhisper languages.English (Nemotron 3) or 40 language-locales (Nemotron 3.5).English and French (1B) or English (2.6B).
Latency (vendor claim)Configurable chunk/min-chunk; papers report a few seconds of latency on GPU (not re-verified).Chunk sizes down to 80 ms per model card; NVIDIA FAQ reports 0.067 s ASR latency at 64 parallel streams for Parakeet CTC 1.1B on 3xH100 (vendor benchmark).Model-defined delay: 0.5 s (1B) or 2.5 s (2.6B) per README.
Key limits
  • Research-grade code: one stream per process, minimal ops tooling.
  • GPU required for the multilingual model.
  • Third-party review notes occasional misplaced punctuation in streaming and inconsistent auto language detection; specify the language when known.
  • Fixed algorithmic delay per model; 2.6B is too slow-to-final for snappy voice agents.
High-severity warnings
  • None
  • None
  • None
ComplianceSelf-hosted.Self-hosted: data stays in your environment.Self-hosted.
Self-hostableYesYesYes
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
TypeSpeech-to-textSpeech-to-textSpeech-to-text
Params B-0.61
VRAM GB---
CPU ok-No-
LicenceMITOpenMDW-1.1CC-BY-4.0
CommercialYesYesYes
Full duplex---
StreamingYesYesYes
Latency ms--500
Languages99402
High warnings000