NVIDIA Nemotron 3: Efficient and Open Intelligence | Summary
24 Dec 2025 | Paper Review Mixture-of-Experts LLM Pretraining Reinforcement Learning Long-Context Language ModelsContents
- Summary
- 1. Introduction
- 2. Features and Technologies
- 3. Key Takeaways
- Contributors
- Appendix
- Brief Thoughts
This article explains the key points of NVIDIA Nemotron 3: Efficient and Open Intelligence.
- 2025-12-24 (arXiv)
- NVIDIA, Blakeman, Aaron, Grattafiori, Aaron, Basant, Aarti, Gupta, Abhibha, Khattar, Abhinav, Renduchintala, Adi, Vavre, Aditya, Shukla, Akanksha, Bercovich, Akhiad, et al.
- NVIDIA
- Paper
Summary
- NVIDIA Nemotron 3: Efficient and Open Intelligence introduces Nano, Super, and Ultra for reasoning, conversation, and agentic workloads. The family uses a hybrid Mamba-Transformer Mixture-of-Experts architecture, supports contexts up to 1M tokens, and combines multi-environment reinforcement learning with inference-time reasoning budget control.
- Nemotron-3-Nano-30B-A3B achieves 3.3× the throughput of Qwen3-30B-A3B at 8k input tokens and 16k output tokens. Figure 2 reports higher Nano scores on the displayed chat, instruction-following, tool-use, coding, and 1M-token retrieval comparisons; AIME25 rankings differ between evaluations with and without tools.
- Super and Ultra incorporate LatentMoE, Multi-Token Prediction (MTP), and NVFP4 training. Separate approximately 8B-active-parameter ablations show improvements in all five LatentMoE task aggregates and all seven MTP benchmark scores, while NVFP4-trained models achieve downstream accuracy comparable to BF16-trained models when evaluated in BF16.
- The white paper presents released Nano results alongside design choices and release plans for Super and Ultra. It provides component ablations but no complete Super or Ultra evaluations, detailed throughput measurement settings, or matched comparison testing the claimed stability and reward-hacking benefits of simultaneous reinforcement learning.
1. Introduction
The introduction addresses the inference cost of long reasoning traces and complex agent interactions. It presents the hybrid architecture as a way to improve the accuracy-throughput trade-off, with contexts up to 1M tokens intended for long code sequences, conversation histories, and documents used in RAG pipelines.
- All three models use diverse reinforcement learning environments and support reasoning budget control.
- LatentMoE, NVFP4 training, and MTP are additional features of Super and Ultra.
- The stated release plan includes model weights, training recipes, pre- and post-training software, and over 10 trillion tokens of datasets; data release is subject to redistribution rights.
2. Features and Technologies
The technical section describes seven components: Hybrid MoE, LatentMoE, MTP, NVFP4 training, long-context preparation, multi-environment reinforcement learning, and reasoning budget control. Together, these address generation cost, expert memory and communication costs, training precision, context extension, and post-training.
2.1. Hybrid MoE
Nemotron 3 predominantly interleaves Mamba-2 and MoE layers, retaining a small number of self-attention layers. Mamba-2 maintains a constant-size recurrent state during generation, while the retained attention layers provide direct information routing across tokens and still require attention caches.
- Figure 1 shows Nano’s layer pattern: a Mamba-2/MoE/Mamba-2/MoE/Mamba-2/Attention/MoE block repeated five times, a Mamba-2/MoE block repeated three times, one Mamba-2/Attention/MoE block, and a final Mamba-2/MoE block repeated four times.
- MoE supplies sparse parameter scaling. The constant-state generation property applies to Mamba-2 layers rather than the complete hybrid model.
The repeated blocks show that most sequence-processing stages pair Mamba-2 with MoE, with attention inserted at selected positions. The diagram makes the limited use of attention explicit while confirming that Nano retains attention layers.
At 8k input tokens and 16k output tokens, Figure 2 reports relative output throughput per GPU of 3.3 for Nemotron-3-Nano-30B-A3B, 1.0 for Qwen3-30B-A3B-Thinking-2507, and 1.5 for GPT-OSS-20B-A4B. Nano scores 86.3 on RULER at 1M tokens versus 77.5 for Qwen; GPT-OSS has no reported result for this comparison.
- Arena-Hard-v2-Avg scores are 67.7, 57.8, and 48.5; IFBench scores are 71.5, 51.0, and 65.0; τ²-Bench scores are 49.0, 47.7, and 47.5, respectively.
- SWE-Bench scores are 38.8, 22.0, and 34.0; LCB v6 scores are 68.2, 66.0, and 61.0, respectively.
- AIME25 scores are 89.1 for Nano, 85.0 for Qwen, and 91.7 for GPT-OSS. Nano’s explicitly tool-assisted score is 99.2. Figure 2 also displays 98.7 for GPT-OSS without specifying that label’s tool-use condition. The white paper refers readers to the separate Nano technical report for evaluation details.
Nano leads most displayed accuracy comparisons and achieves relative throughput of 3.3 versus Qwen's 1.0 at ISL/OSL 8k/16k. Nano's AIME25 score rises from 89.1 to the explicitly tool-assisted 99.2. GPT-OSS has labels of 91.7 and 98.7; this figure does not specify the second label's tool-use condition. GPT-OSS has no reported RULER result at 1M tokens.
2.2. LatentMoE: Hardware-Aware Expert Design for Improved Accuracy per Byte
LatentMoE targets expert-weight memory traffic in latency-focused deployments processing tens to hundreds of tokens and all-to-all communication in throughput-focused deployments processing thousands of tokens. Expert matrices have size d × m, communication volume scales linearly with K and d, and the effective nonlinear budget is described as roughly proportional to K × m.
- Here d is the model hidden dimension, m is the expert FFN intermediate dimension, and K is the number of active experts per token.
- Reducing the routed input dimension can reduce both weight-loading traffic and communication payloads; reducing m does not directly reduce token-dispatch communication.
Tokens are projected from hidden dimension d into a smaller latent dimension ℓ < d for routed expert computation, then projected back to d. Per-expert weight loads and dispatch payloads decrease by a factor of d/ℓ, typically about 4×; the proposed scaling uses these savings to increase expert counts through N′ = N · d/ℓ and K′ = K · d/ℓ.
- The routing gate, shared expert, and non-expert layers remain in the original hidden dimension d.
- The experimental Standard MoE uses d = 4096, 128 total experts, and 6 active experts; LatentMoE uses ℓ = 1024, 512 total experts, and 22 active experts. The experimental active-expert count differs from the illustrative fourfold scaling.
The down-projection and up-projection surround routed expert computation and communication, allowing more experts to process narrower representations. The shared expert and router remain outside this compression, matching the text's distinction between routed and non-routed computation.
Table 1 compares Standard MoE with 8.09B active and 72.6B total parameters against LatentMoE with 8.02B active and 72.8B total parameters, both trained for 1 trillion tokens using identical hyperparameters. LatentMoE improves MMLU-Pro from 48.30 to 52.87, MMLU from 70.10 to 72.11, Code from 51.95 to 55.14, Math from 78.32 to 80.19, and Commonsense Understanding from 81.73 to 82.10.
- Code averages HumanEval, HumanEval+, MBPP, and MBPP+; Math averages GSM8K CoT and MATH-500.
- Commonsense Understanding averages RACE, ARC-Challenge, HellaSwag, and Winogrande.
- The comparison demonstrates an accuracy benefit at closely matched parameter counts but provides no measured latency or throughput results.
LatentMoE exceeds Standard MoE in all five columns after 1 trillion training tokens with identical hyperparameters and closely matched parameter counts. The largest absolute gain is MMLU-Pro, from 48.30 to 52.87; the smallest is Commonsense Understanding, from 81.73 to 82.10. The table contains no measured inference costs.
2.3. Multi-Token Prediction (MTP)
MTP predicts multiple future tokens, providing auxiliary training signals and draft tokens for speculative decoding without a separate draft model. In an 8B-active-parameter Transformer MoE base model trained on 1T tokens, Table 2 shows gains across general knowledge, code, commonsense understanding, reading comprehension, and mathematics; the authors report roughly 2.4% average improvement.
- MMLU (5-shot, acc) improves from 70.06 to 71.26; MMLU-Pro (5-shot, CoT EM) from 45.05 to 47.84; MBPP-Sanitized (3-shot) from 65.58 to 66.89.
- ARC-Challenge (25-shot, acc_norm) improves from 86.43 to 88.05; WinoGrande (0-shot, acc) from 74.59 to 75.45; RACE (0-shot, acc) from 84.02 to 85.36; GSM8K (8-shot, acc) from 82.49 to 84.46.
- A lightweight MTP module achieves around 97% acceptance on the first two predicted tokens in an 8B-active MoE ablation. The section provides no measured end-to-end decoding speedup.
Adding MTP improves every reported benchmark under the listed evaluation conditions. Gains range from 0.86 percentage points on WinoGrande to 2.79 percentage points on MMLU-Pro, demonstrating benefits across the seven tasks in this Transformer MoE ablation.
2.4. NVFP4 Training
The authors report stable hybrid Mamba-MoE pretraining for up to 25T tokens using native NVFP4 GEMMs through the cuBLAS backend for Transformer Engine. Quantizing weights, activations, and gradients enables NVFP4 operations in fprop, dgrad, and wgrad, subject to the recipe’s precision exceptions; the cited 3× FP4-versus-FP8 throughput on GB300 is a hardware peak comparison.
- NVFP4 uses 16-element micro-block scaling, E4M3 block scaling factors, a second-level FP32 global scale, and E2M1 elements.
- The recipe includes two-dimensional weight block scaling, Random Hadamard Transforms on inputs to wgrad, and stochastic rounding on gradients.
- The final 15% of the network remains in high precision. Latent projections and MTP layers remain in BF16.
QKV and attention projections remain in BF16, and Mamba output projections remain in MXFP8 because NVFP4 quantization can flush up to 40% of their values to zero on Nano. Attention layers use GQA with only 2 KV heads.
- Figure 4 compares the mixed-precision recipe with an ablation that quantizes sensitive layers starting from a 500B-token Nano checkpoint; the ablation produces larger loss gaps.
- The authors report relative loss differences below 1% for Nano and below 0.6% for a larger 8B-active MoE model compared with BF16. Figure 4 includes transient excursions, so these thresholds do not describe every checkpoint.
- Figure 5 follows an 8B-active model through 1T training tokens and shows comparable downstream trajectories. All displayed evaluations use BF16, establishing evidence about training quality rather than NVFP4 inference accuracy.
The later loss trajectories show a smaller BF16 gap for A8B than for A3B, although the training curves include transient excursions. Quantizing Nano's sensitive layers after the 500B-token checkpoint increases the gap, supporting the recipe's use of higher precision for selected projections.
Across eight downstream tasks, NVFP4-trained and BF16-trained models follow broadly similar accuracy trajectories through 1T tokens, with differences in both directions. Evaluation uses BF16, so the comparison tests how training precision affects learned model quality.
2.5. Long Context
Nemotron 3 omits RoPE from attention layers because Mamba layers provide implicit positional information. Nano receives continued pre-training at 512k sequence length, supervised fine-tuning at 256k, and reinforcement learning with long-context inputs up to 32k tokens; all three stages use synthetic data for retrieval, multi-hop reasoning, and multi-document aggregation.
- The authors report that continued pre-training did not require a staged increase from 8k to 512k.
- The supported 1M-token context exceeds the 512k continued-pretraining sequence length, so performance at 1M depends on length extrapolation.
Table 3 compares two base models trained up to 512k sequence length. Nemotron-3-Nano-30B-A3B-Base scores 54.19 on RULER at 1M tokens versus 23.43 for Nemotron-Nano-12B-v2-Base, but scores lower at 128k, 256k, and 512k; the comparison supports more robust extrapolation beyond training length.
- The dense hybrid scores 85.13, 79.85, and 75.12 at 128k, 256k, and 512k; the MoE hybrid scores 74.48, 71.67, and 66.02.
- Figure 6 shows declining cumulative average NLL on held-out repository-level code sequences larger than 1 million tokens, with a power-law fit reporting R² = 0.883.
- The NLL trend is consistent with useful prediction over long code contexts, but cumulative averaging does not isolate the contribution of distant tokens or establish agent-task performance at 1M tokens.
The dense hybrid leads at 128k, 256k, and 512k, but drops from 75.12 at 512k to 23.43 at 1M. The MoE hybrid declines from 66.02 to 54.19 over the same interval, showing more robust extrapolation beyond their shared maximum training length of 512k.
Cumulative average NLL decreases as the evaluated code sequence extends toward 1M tokens, and the displayed power-law fit has R² = 0.883. This trend is consistent with useful prediction over long code contexts, but it does not isolate the contribution of particular distant tokens.
2.6. Multi-environment Reinforcement Learning Post-training
Post-training jointly optimizes mathematical and scientific reasoning, competitive coding, instruction following, software engineering, search, chat, tool use, and long-context tasks. The authors report greater stability and less reward hacking than in earlier staged approaches, while Figure 7 shows benchmark trajectories within a Nano RL run.
- Several curves improve clearly, including GPQA, LiveCodeBench, and MMLU Pro; SciCode and τ²-Bench fluctuate substantially.
- The figure documents joint optimization but includes no matched staged-training comparison or quantitative reward-hacking measurement.
The training system decouples inference from optimization through asynchronous RL and uses MTP to accelerate rollout generation. GRPO with masked importance sampling accounts for discrepancies between training and rollout policies.
- The paper states that the post-training software stack is open-sourced under Apache 2.0.
- NeMo-RL implements scalable training at https://github.com/NVIDIA-NeMo/RL, and NeMo-Gym supplies environments at https://github.com/NVIDIA-NeMo/Gym.
GPQA, LiveCodeBench, MMLU Pro, and several other benchmarks improve over the RL run. SciCode fluctuates after an early rise, and τ²-Bench varies substantially without a sustained upward trend. The plot documents joint optimization across capabilities rather than uniform monotonic improvement.
2.7. Granular Reasoning Budget Control at Inference Time
Users can cap the tokens in a thinking trace and append </think> when the budget is reached, allowing the model to answer from the partial trace. Figure 8 plots accuracy against average generated tokens per query on AIME25, GPQA, LiveCodeBench, and MMLU Pro.
- Accuracy generally increases with generation length, with task-specific saturation and small downturns near some of the largest displayed budgets.
- The horizontal axis measures average generated tokens per query rather than wall-clock latency or deployment cost.
Longer average generations generally improve accuracy, with different useful ranges across tasks. AIME25 continues improving toward 32K generated tokens, while GPQA, LiveCodeBench, and MMLU Pro flatten or show small declines near their largest displayed lengths, supporting task-specific budget selection.
3. Key Takeaways
The concluding section reiterates the shared hybrid architecture, multi-environment RL, reasoning budget control, and support for contexts up to 1M tokens. It identifies LatentMoE, NVFP4 training, and MTP as additional Super and Ultra features, and states that Nano accompanies the white paper while Super and Ultra releases will follow in the upcoming months.
- The planned release covers model weights, training software, recipes, and most training data, subject to redistribution rights.
- The reported evidence consists of Nano comparisons and component ablations; complete Super and Ultra evaluations are not presented.
Contributors
The Contributors section credits Pretraining Data, Architecture, Pretraining Software, Pretraining, NVFP4, Long Context, Posttraining Software, and Posttraining. It continues with Evaluation, Safety and Release, Infrastructure, Quantization, Inference, Compression, Deployment, Legal and Compliance, Marketing, Project Management, Product, and Leadership.
Appendix
- The supplied paper has no source appendix. Contributors and References follow the main technical sections, which direct readers to the separate Nemotron Nano 3 technical report for detailed Nano evaluation settings.
Brief Thoughts
LatentMoE reduces the dimension governing both expert-weight traffic and dispatch payloads, then uses the savings to increase expert counts. Table 1 supports an accuracy benefit at closely matched active and total parameter counts, but measured latency and communication costs are needed to test the proposed systems benefit.
The precision ablation identifies sensitive layers and tests a mixed-precision recipe against BF16. The displayed downstream comparison extends to 1T training tokens; neither that study nor the cited hardware peak throughput establishes downstream quality or realized training speed across the complete reported up-to-25T-token runs.
The MoE base model retains more RULER accuracy when extrapolating from 512k to 1M tokens, despite trailing the dense comparator at the shorter evaluated lengths. The joint RL curves show gains on several benchmarks, but their variability and the absence of a staged-training control leave the claimed stability and reward-hacking advantages untested by the displayed comparison.