Particle.news

Nvidia Releases 30B Nemotron 3.5 Lightning and Open NeMo Switchyard

Nvidia says the tools will cut inference costs by routing tasks to smaller open models for faster agent deployments.

Overview

  • Nvidia, which launched the software Tuesday, made Nemotron 3.5 Lightning and NeMo Switchyard publicly available for download and integration.
  • Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model that Nvidia says can run on a single GPU and generate tokens up to four times faster than peers.
  • NeMo Switchyard is an open-source routing library written in Rust that lets agents send each step to the best-fit model based on latency, cost or quality.
  • Early partner tests and Nvidia internal benchmarks report large cost and runtime gains from routing, with LangChain noting a 74% cost cut at a 6% accuracy tradeoff and Ramp reporting 58% lower cost and 33% faster runtime.
  • Independent indices still rank Nemotron behind rivals such as Google’s Gemma 4, and the release sharpens debate over distillation, post-training tradeoffs for specialization, and demand for GPUs for local and enterprise deployments.