NvidiaNvidia·💬 Text Generation·New

NVIDIA Nemotron 3.5 Lightning 30B

ReasoningCodeFunction CallingWeb Searchfp4private
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
NVIDIA Nemotron 3.5 Lightning 30B — TLDR
  • 🧠 Hybrid Mamba-2 + Mixture-of-Experts model: 30B total, 3B active parameters.
  • 📏 Context length up to 1M tokens, per NVIDIA's model card.
  • ⚡ Ships with speculative decoding methods and Multi-Token Prediction layers for faster generation.
  • 🔧 LatentMoE routing projects tokens into a smaller latent space before expert compute.
  • 🎯 Aimed at high-volume, low-latency execution inside always-on agentic workflows.
  • 📚 Pre-trained on over 20T tokens using an NVFP4 recipe.
  • 🌐 English plus 19 other spoken languages and 43 programming languages.
  • 🏢 NVIDIA-released open weights, marked ready for commercial use.
💰 Pricing
$0.100 / $0.250
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Aug 12, 2026
5 days ago
Provider

Nvidia Corporation is an American technology company founded in 1993 by Jensen Huang, Chris Malachowsky, and Curtis Priem, headquartered in Santa Clara, California. Long recognized as the dominant force in graphics processing units, Nvidia has expanded into a…

Read full profile →
5 models on Venice
3 text · 1 embedding · 1 asr
Since Oct 10, 2025

About this model

NVIDIA Nemotron 3.5 Lightning 30B is the speed-tier entry in NVIDIA's Nemotron line, positioned for the "execution layer" of long-running agents: tool calls, result validation, and subagent delegation rather than deep frontier reasoning. It is a hybrid Mixture-of-Experts model with 30B total and 3B active parameters, interleaving Mamba-2 and MoE layers with select attention layers, and supports context lengths up to 1M tokens.

Within the catalog it sits in the same size class as NVIDIA Nemotron 3 Nano 30B, while NVIDIA Nemotron 3 Ultra is the much larger Nemotron tier. The generational changes are architectural: NVIDIA describes a LatentMoE design in which tokens are projected into a smaller latent dimension for expert routing and computation, which the company says improves accuracy per byte, plus Multi-Token Prediction layers that provide richer training signal. The release also bundles several speculative decoding methods specifically for faster text generation.

Efficiency extends to deployment. Pre-training used more than 20T tokens with an NVFP4 recipe, and NVIDIA publishes both a BF16 checkpoint and an NVFP4 post-training-quantized variant using W4A16 on routed and shared experts with FP8 dynamic scales on Mamba projections and KV cache, plus a recipe tuned for DGX Spark.

It is intended as a general-purpose reasoning and chat model for developers building AI agents, chatbots, and RAG systems, covering English, 19 additional spoken languages, and 43 programming languages.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 4d ago