Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| WhisperLiveKit Open source (QuentinFuxa) | NVIDIA Nemotron ASR Streaming / Parakeet (Riva, NIM, NeMo) NVIDIA | Kyutai STT (Delayed Streams Modeling) Kyutai | |
|---|---|---|---|
| Category | Open models | Open models | Open models |
| Status | GA | GA | GA |
| Est. per minute | $0 | $0 | $0 |
| How that was worked out | Self-host compute only. | Self-host compute only; cost depends on GPU price and streams per GPU. | Self-host compute only. |
| Pricing model | free | free | free |
| Free tier | Open source. | Open weights; hosted API trial on build.nvidia.com. | Open weights. |
| Connects by | WebSocket, HTTP (web UI) | gRPC (Riva/NIM), WebSocket (Riva realtime client), Python (NeMo) | WebSocket (moshi-server), Python (PyTorch, MLX) |
| Audio in | Browser mic via built-in web page or WebSocket clients. | 16 kHz mono PCM typical for Riva streaming (not re-verified per model). | Handled by the provided scripts (mic or file); server streams PCM over WebSocket. |
| Audio out | n/a | n/a | n/a |
| Languages | Whisper languages. | English (Nemotron 3) or 40 language-locales (Nemotron 3.5). | English and French (1B) or English (2.6B). |
| Latency (vendor claim) | Policy and model dependent. | Chunk sizes down to 80 ms per model card; NVIDIA FAQ reports 0.067 s ASR latency at 64 parallel streams for Parakeet CTC 1.1B on 3xH100 (vendor benchmark). | Model-defined delay: 0.5 s (1B) or 2.5 s (2.6B) per README. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Self-hosted. | Self-hosted: data stays in your environment. | Self-hosted. |
| Self-hostable | Yes | Yes | Yes |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Type | Speech-to-text | Speech-to-text | Speech-to-text |
| Params B | - | 0.6 | 1 |
| VRAM GB | - | - | - |
| CPU ok | - | No | - |
| Licence | Apache-2.0 | OpenMDW-1.1 | CC-BY-4.0 |
| Commercial | Yes | Yes | Yes |
| Full duplex | - | - | - |
| Streaming | Yes | Yes | Yes |
| Latency ms | - | - | 500 |
| Languages | 99 | 40 | 2 |
| High warnings | 0 | 0 | 0 |