NVIDIA Nemotron 3.5 Lightning 30B
About this model
NVIDIA Nemotron 3.5 Lightning 30B is the speed-tier entry in NVIDIA's Nemotron line, positioned for the "execution layer" of long-running agents: tool calls, result validation, and subagent delegation rather than deep frontier reasoning. It is a hybrid Mixture-of-Experts model with 30B total and 3B active parameters, interleaving Mamba-2 and MoE layers with select attention layers, and supports context lengths up to 1M tokens.
Within the catalog it sits in the same size class as NVIDIA Nemotron 3 Nano 30B, while NVIDIA Nemotron 3 Ultra is the much larger Nemotron tier. The generational changes are architectural: NVIDIA describes a LatentMoE design in which tokens are projected into a smaller latent dimension for expert routing and computation, which the company says improves accuracy per byte, plus Multi-Token Prediction layers that provide richer training signal. The release also bundles several speculative decoding methods specifically for faster text generation.
Efficiency extends to deployment. Pre-training used more than 20T tokens with an NVFP4 recipe, and NVIDIA publishes both a BF16 checkpoint and an NVFP4 post-training-quantized variant using W4A16 on routed and shared experts with FP8 dynamic scales on Mamba projections and KV cache, plus a recipe tuned for DGX Spark.
It is intended as a general-purpose reasoning and chat model for developers building AI agents, chatbots, and RAG systems, covering English, 19 additional spoken languages, and 43 programming languages.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 4d ago