Compare realtime APIs
Pick up to four APIs from any category and see them side by side.
| Kyutai STT (Delayed Streams Modeling) Kyutai | NVIDIA Nemotron ASR Streaming / Parakeet (Riva, NIM, NeMo) NVIDIA | Moonshine Voice (Moonshine v2 streaming) Moonshine AI (Useful Sensors) | |
|---|---|---|---|
| Category | Open models | Open models | Open models |
| Status | GA | GA | GA |
| Est. per minute | $0 | $0 | $0 |
| How that was worked out | Self-host compute only. | Self-host compute only; cost depends on GPU price and streams per GPU. | Runs on the user's device. |
| Pricing model | free | free | free |
| Free tier | Open weights. | Open weights; hosted API trial on build.nvidia.com. | Free and open source. |
| Connects by | WebSocket (moshi-server), Python (PyTorch, MLX) | gRPC (Riva/NIM), WebSocket (Riva realtime client), Python (NeMo) | Local library (Python, JS/WASM, iOS, Android, desktop) |
| Audio in | Handled by the provided scripts (mic or file); server streams PCM over WebSocket. | 16 kHz mono PCM typical for Riva streaming (not re-verified per model). | Microphone or PCM via the library. |
| Audio out | n/a | n/a | n/a |
| Languages | English and French (1B) or English (2.6B). | English (Nemotron 3) or 40 language-locales (Nemotron 3.5). | English plus additional languages (list on docs site; not captured). |
| Latency (vendor claim) | Model-defined delay: 0.5 s (1B) or 2.5 s (2.6B) per README. | Chunk sizes down to 80 ms per model card; NVIDIA FAQ reports 0.067 s ASR latency at 64 parallel streams for Parakeet CTC 1.1B on 3xH100 (vendor benchmark). | Paper: bounded time-to-first-token independent of utterance length; third-party claim of sub-200 ms on edge devices is unverified. |
| Key limits |
|
|
|
| High-severity warnings |
|
|
|
| Compliance | Self-hosted. | Self-hosted: data stays in your environment. | On-device processing. |
| Self-hostable | Yes | Yes | Yes |
| Last checked | 2026-10-10 | 2026-10-10 | 2026-10-10 |
| Key numbers and features | |||
| Type | Speech-to-text | Speech-to-text | Speech-to-text |
| Params B | 1 | 0.6 | - |
| VRAM GB | - | - | - |
| CPU ok | - | No | Yes |
| Licence | CC-BY-4.0 | OpenMDW-1.1 | MIT |
| Commercial | Yes | Yes | Yes |
| Full duplex | - | - | - |
| Streaming | Yes | Yes | Yes |
| Latency ms | 500 | - | - |
| Languages | 2 | 40 | - |
| High warnings | 0 | 0 | 0 |