Back to events
Model releaseModerate global significanceConfirmed confidence

NVIDIA releases Nemotron 3.5 ASR Streaming 0.6B

NVIDIA releases a 600M multilingual, cache-aware streaming ASR model.

Event details

The model processes continuous audio in non-overlapping cached chunks and supports multiple streaming-latency settings, punctuation, capitalization and automatic language detection.

Why it matters

Open weights and a cache-aware architecture improve low-latency concurrency and local deployment efficiency for multilingual streaming ASR.

68/100Global significance score. Regional effects are recorded only when the evidence supports a meaningful difference.

Access notes

The weights are downloadable and support local execution with NeMo, Transformers and NeMo-Speech.cpp.