GA NVIDIA

NVIDIA Nemotron ASR Streaming / Parakeet (Riva, NIM, NeMo)

Open-weight cache-aware FastConformer-RNNT streaming ASR models (600M) that you serve with Riva/NIM containers or NeMo. Nemotron 3.5 ASR adds 40 language-locales in one checkpoint.

Est. per minute$0

Overview

Best for: High-volume, low-latency English or multilingual streaming on your own NVIDIA GPUs, especially for voice agents.

At a glance

TypeSpeech-to-text
Params B0.6
CPU okNo
LicenceOpenMDW-1.1
CommercialYes
StreamingYes
Languages40

Values for Nemotron 3.5 ASR (40 language-locales); the English Nemotron 3 model uses the NVIDIA Open Model License. Chunks down to 80 ms; no end-to-end latency figure for this model. Hosted trial on build.nvidia.com.

Audio in

16 kHz mono PCM typical for Riva streaming (not re-verified per model).

Audio out

n/a

Languages

English (Nemotron 3) or 40 language-locales (Nemotron 3.5).

Latency

Chunk sizes down to 80 ms per model card; NVIDIA FAQ reports 0.067 s ASR latency at 64 parallel streams for Parakeet CTC 1.1B on 3xH100 (vendor benchmark).

Regions

Self-hosted anywhere.

Compliance

Self-hosted: data stays in your environment.

Hardware

NVIDIA GPU (H100/L40S/A10-class typical); many concurrent streams per GPU via batching.

Licence

Nemotron 3 EN: NVIDIA Open Model License; Nemotron 3.5: OpenMDW-1.1; Parakeet TDT 0.6B v3: CC-BY-4.0; parakeet_realtime_eou: NVIDIA licence (other).

Features

  • cache-aware streaming (no re-processing overlap)
  • runtime latency/accuracy trade-off
  • punctuation and capitalization
  • batching many streams per GPU
  • fine-tuning with NeMo

Pricing

WhatPriceUnitNotes
Open weights$0You pay for GPUs. NVIDIA AI Enterprise licence may apply to production Riva/NIM support.
Hosted trial (build.nvidia.com)free trial creditsHosted nemotron-asr-streaming endpoint for evaluation.
How the per-minute estimate was worked out

Self-host compute only; cost depends on GPU price and streams per GPU.

Free tier: Open weights; hosted API trial on build.nvidia.com.

Source: huggingface.co

Setup

  1. Pick a model on Hugging Face and read its licence (Nemotron EN: NVIDIA Open Model License; Nemotron 3.5: OpenMDW-1.1).
  2. Production: run the Riva or NIM ASR container (NGC nemotron-asr-streaming) on an NVIDIA GPU.
  3. Use nvidia-riva python-clients scripts (transcribe_mic.py, realtime_asr_client.py) to stream audio over gRPC.
  4. Research/fine-tuning: use NeMo cache-aware streaming inference scripts.

Endpoint

localhost:50051 gRPC (Riva/NIM default) or your service URL

Authentication

None by default on self-hosted Riva; NGC API key to pull containers; NVIDIA API key for hosted trial.

Quick start bash

# After starting a Riva/NIM ASR server on localhost:50051
pip install nvidia-riva-client
git clone https://github.com/nvidia-riva/python-clients && cd python-clients
python scripts/asr/transcribe_mic.py --server localhost:50051 --language-code en-US
# or stream a file at real-time pace:
python scripts/asr/transcribe_file.py --server localhost:50051 --input-file audio_16k_mono.wav

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Licences differ per model

nemotron-speech-streaming-en uses the NVIDIA Open Model License, Nemotron 3.5 uses OpenMDW-1.1, Parakeet TDT v3 is CC-BY-4.0, and Riva/NIM containers fall under NVIDIA licence terms. Read each before shipping.

Multilingual model needs a GPU

Claims that the 40-language model runs on CPU refer to separate work on the English model; plan for NVIDIA GPUs.

You own scaling and uptime

Self-hosting means managing GPU capacity, batching, health checks and upgrades; throughput claims are vendor benchmarks on H100s.

Conflicting release dates

Hugging Face shows Nemotron 3.5 released 2026-06-04 while an OpenRouter listing shows 2026-08-13 for a dated snapshot; pin the exact revision.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • GPU required for the multilingual model.
  • Third-party review notes occasional misplaced punctuation in streaming and inconsistent auto language detection; specify the language when known.

Models and products

NameStatusNotes
nvidia/nemotron-speech-streaming-en-0.6b (Nemotron 3 ASR)GA (open weights)English, punctuation and capitalization; latency chosen at inference via att_context_size in 80 ms frames.
nvidia/nemotron-3.5-asr-streaming-0.6bGA (open weights, HF 2026-06-04)40 language-locales with language-ID prompt conditioning; GPU required.
Parakeet CTC / RNNT / TDT family (e.g. parakeet-tdt-0.6b-v3)GAOlder Parakeet models; TDT v3 is CC-BY-4.0.
nvidia/parakeet_realtime_eou_120m-v1Released 2025Small realtime model with end-of-utterance detection.

Docs and sources

Docs

Sources used

Not fully verified

Exact CLI flags of the Riva client scripts, Riva/NIM licence terms for production, and per-GPU stream counts.

Similar open models APIs

Spotted a wrong price or a dead link?