Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding | Summary
07 Jul 2026 | Paper Review Diffusion Language Models Speculative Decoding Autoregressive Language Models Parallel DecodingContents
- Summary
- 2. Tri-Mode LM Training
- 3. Tri-Mode LM Inference
- 4. Speed-of-Light Analysis
- 5. Nemotron-Labs-Diffusion Family
- 6. Evaluation and Analysis
- 8. Insights and Future Directions
- Appendix
- Brief Thoughts
This article explains the key points of Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding.
- 2026-07-07 (arXiv)
- Fu, Yonggan, Whalen, Lexington, Garg, Abhinav, Wu, Chengyue, Khadkevich, Maksim, Oswald, Nicolai, Xie, Enze, Egert, Daniel, Sreenivas, Sharath Turuvekere, Diao, Shizhe, et al.
- NVIDIA, Georgia Tech, HKU, University of Chicago, MIT
- Paper
Summary
- Autoregressive (AR) decoding generates tokens sequentially; diffusion decoding can predict positions in parallel but faces accuracy and efficiency trade-offs. Nemotron-Labs-Diffusion jointly trains AR next-token prediction and block-wise diffusion denoising in one backbone, supporting AR, diffusion, and self-speculation decoding without separate prediction heads.
- In self-speculation, diffusion drafts tokens and the same model’s AR pathway verifies them. On 10 instruct tasks, the 8B model averages 63.61 accuracy in AR mode and 63.18 in diffusion mode; LoRA-enhanced linear self-speculation averages 62.81 at 5.99 tokens per forward (TPF), compared with Qwen3-8B’s 62.75 at 1.00 TPF.
- Oracle-guided diffusion reaches 6.02 TPF on SPEED-Bench versus 3.41 for linear self-speculation, but the oracle result is not a deployed sampler measurement. The family includes 3B/8B/14B base and instruct models and an 8B vision-language model; SGLang tests separately measure low-concurrency throughput on GB200, RTX Pro 6000, and DGX Spark.
The panels relate decoding mode to operating conditions: AR is presented for higher concurrency, while self-speculation improves single-stream throughput in the reported measurements. The GB200 curves also show why a batch-size-1 advantage cannot be applied unchanged to high-concurrency serving.
The comparison displays overall, coding, and math accuracy alongside TPF for the 8B instruct model and baselines. It makes the task scores visible next to the gains in tokens produced per forward.
2. Tri-Mode LM Training
Training combines a clean causal stream with a corrupted stream so the shared backbone receives AR and in-block denoising supervision. The recipe uses an AR-only initialization followed by joint training and proposes global token-level loss averaging for variable masking ratios.
- The attention design follows earlier joint AR–diffusion work [14]; this paper additionally studies tri-mode inference, post-training enhancements, a model family, and an oracle-guided diffusion analysis.
2.1. Training Objectives
The AR objective predicts each token from its prefix; block-wise diffusion denoises corrupted positions conditioned on clean preceding blocks. The joint objective is L(θ) = L_AR(θ) + αL_diff(θ): Stage 1 sets α = 0, and Stage 2 uses α = 0.3.
- The authors argue that per-sequence averaging can overweight examples with few masked tokens because diffusion loss includes a 1/t factor. Their displayed Eqs. 4–5, however, both divide fixed-length token sums by NL; without an explicit active-token normalization, those equations are algebraically equal.
2.2. Attention Pattern
Noisy tokens attend bidirectionally within a block and to the clean prefix, while the clean stream uses strictly causal attention. The clean-to-clean restriction permits AR and diffusion losses in one forward–backward pass without exposing future tokens to AR predictions.
The mask separates noisy-to-noisy, noisy-to-clean-prefix, and causal clean-to-clean attention. The last restriction prevents future-token leakage into AR supervision computed during the same training pass.
2.3. Ablation Study on Training Techniques
In the progressive ablation starting from Ministral3-8B, diffusion-mode mean accuracy rises from 54.23 with block-wise attention to 56.35 after global loss averaging, 57.06 after DP-rank varying masking ratios, 62.80 after two-stage training, and 70.28 after adding AR loss. These are cumulative comparisons, not isolated factorial effects; the two-stage setting also introduces a 1T-token AR initialization ahead of the reported 25B-token continuous-pretraining comparison.
The progressive rows raise diffusion-mode mean accuracy from 54.23 to 70.28. Because additions are cumulative and the two-stage setting includes further AR initialization, the increments are not isolated effects under an identical full training budget.
2.4. Mutual Impact of AR/Diffusion Losses
In the 25B-token coefficient sweep, α = 0.3 gives the highest reported mean accuracies in both diffusion and AR modes, 69.77 and 70.62. Matched-token comparisons show AR means changing from 72.50 to 72.64 for base models and 63.18 to 63.61 for instruct models when diffusion loss is added, although instruct IFEval falls from 71.66 to 68.65.
- Figure 4 marks the diffusion loss for α = 0 at approximately 12.9, off its plotted scale; training-loss curves alone do not establish the benchmark-optimal coefficient.
Without AR supervision, the plotted AR loss rises sharply; the α = 0 diffusion loss is marked at approximately 12.9, outside the diffusion panel's scale. The benchmark sweep, rather than these loss curves alone, supports choosing α = 0.3.
The reported AR and diffusion means both peak at α = 0.3 in this six-benchmark sweep, at 70.62 and 69.77. Increasing diffusion weight beyond that value is not uniformly helpful.
Adding diffusion loss changes the mean AR scores only slightly in the matched-token base and instruct comparisons, but individual tasks move in both directions. The instruct IFEval decrease from 71.66 to 68.65 limits a blanket claim of improvement.
3. Tri-Mode LM Inference
Inference supports standard causal AR generation, iterative block-wise diffusion, and diffusion drafting with AR verification. A quadratic self-speculation variant combines drafting and verification in one structured-attention pass using a quadratic layout; despite higher TPF, its reported real-device efficiency trails linear self-speculation with the evaluated kernels.
3.2. Mode 2: Block-wise Diffusion Denoising
Diffusion starts each block with masks, iteratively commits positions selected by a confidence threshold, and refreshes the block’s KV cache when complete. A separately trained sampler estimates whether a position’s current top-1 token matches the token eventually committed on a reference denoising trajectory; thresholding its predictions sets an accuracy–TPF trade-off.
3.3. Mode 3: Self-Speculation Decoding
Linear self-speculation appends k masks to a verified prefix, drafts their tokens in parallel, and uses a second causal forward pass to accept the longest qualifying prefix. The first rejected position supplies an additional AR token; both passes can reuse the verified prefix’s KV cache, but later drafts are discarded after a rejection.
- A diffusion-path-only LoRA adapter on attention o_proj (rank 128, α = 512; approximately 36M trainable parameters) aligns drafts with a frozen AR verifier. Its LK-hybrid and cross-entropy losses use the accepted prefix plus the first rejected position, excluding logits conditioned on a continuation the inference loop discards.
The self-speculation panel depicts parallel diffusion drafting, causal verification, rejection, and caching of the accepted prefix. Unlike the separate AR and diffusion modes, linear self-speculation needs two model forwards per cycle.
4. Speed-of-Light Analysis
The speed-of-light (SOL) analysis estimates how many positions diffusion could safely commit per forward while reproducing its own serial-denoising output. It uses no AR verifier, and its reference sequence is not independent ground truth; SOL TPF is an oracle-guided comparison rather than measured serving throughput.
4.1. Diffusion SOL Construction
Serial denoising commits the highest-confidence position once per forward to define a target for a block of length B. Recursive dynamic compaction searches confidence-ranked matching positions for a large subset whose simulated continuation reaches that target, treating candidates that exceed a 5000-forward-pass-per-block simulation budget as unsafe; committing every current match need not preserve the target.
In the example, committing several individually matching tokens can make the remaining predictions diverge from the serial target. The search instead checks whether a selected commit set permits a continuation to that target.
4.2. SOL Evaluation on SPEED-Bench
Across 713 SPEED-Bench samples in 11 categories, SOL acceptance rises from 2.89 at block length 4 to 7.60 at length 32; at length 32 it is 8.01 on STEM, 3.49 on roleplay, and 11.26 on multilingual. Mean accuracy on the separate 10-task instruct suite peaks at block length 8 (65.43), rather than 32 (61.81).
- At length 32, SOL and linear self-speculation have acceptance rates of 7.60 and 6.82, but real TPF of 6.02 and 3.41. Their correctness references differ, and SOL’s 76.5% TPF advantage is not a realized sampler speedup.
The panels distinguish accepted tokens per step from tokens per model forward. Overall acceptance is 7.60 for SOL versus 6.82 for linear SS, while real TPF is 6.02 versus 3.41; linear SS pays for separate draft and verification passes.
Longer blocks increase oracle-guided acceptance, with STEM reaching 8.01 at block length 32, but mean accuracy on the separate instruct benchmark suite peaks at length 8. The category variation also shows that the measured potential parallelism depends on content.
5. Nemotron-Labs-Diffusion Family
The 3B/8B/14B base models start from pretrained Ministral3 models, receive 1T tokens of AR training, then 300B tokens of joint training at α = 0.3 with sequence length 4096 on 256 NVIDIA H100 GPUs. Instruct models receive joint-objective supervised fine-tuning on 45B tokens at sequence length 16k; prompt tokens remain unmasked, and loss is computed only on answer tokens.
5.3. Vision-Language Models
The 8B vision-language model combines the diffusion-aware instruct LM backbone with the vision encoder and two-layer MLP projector from Ministral3-8B-Instruct-2512, then fine-tunes the components together. Its asymmetric dual stream omits unmasked vision tokens from the noisy half but retains them in the clean half, changing total input length from 2L to L_text + L, where L_text = L − N_vis.
6. Evaluation and Analysis
The evaluations cover non-thinking instruct models, base models, and VLMs; TPF is reported separately from device throughput. At 3B and 14B, LoRA-tuned linear self-speculation attains 4.36 and 5.96 TPF with mean accuracies of 55.00 and 66.36; the 8B base model reaches 72.36 at 4.67 TPF, versus Qwen3-8B base’s 71.58 at 1.00 TPF.
- The 3B instruct comparison includes Qwen3-4B rather than a parameter-matched 3B comparator. Increasing TPF with model scale is observed here, but the tables do not isolate the proposed future-prediction explanation.
6.1. Benchmark Instruct Models
Across GPQA, IFEval, MMLU, three coding tasks, and four math tasks, the 8B instruct model averages 63.61/63.18/62.81/64.04 accuracy in AR/diffusion/linear SS/quadratic SS modes, at 1.00/2.57/5.99/6.38 TPF. At block length 32, the trained sampler improves the plotted diffusion accuracy–TPF frontier, including 1.3× TPF at matched accuracy or 10.6 percentage points of accuracy at matched TPF; LoRA raises 8B linear-SS TPF from 4.52 to 5.99 with similar mean accuracy.
- AR baselines and the proposed model use NeMo-Skills [26], whereas diffusion baselines use their original evaluation pipelines; cross-model results therefore do not share one evaluation implementation.
Linear SS reaches 5.99 TPF with 62.81 mean accuracy across 10 instruct tasks, near the proposed model's AR and diffusion means. Individual comparisons vary substantially; its IFEval result remains below Qwen3-8B's.
LoRA increases linear-SS TPF at each of the 3B, 8B, and 14B scales with little change to the displayed mean accuracy. The 8B comparison rises from 4.52 to 5.99 TPF, directly testing drafter tuning.
For block length 32, the trained sampler improves the displayed accuracy–TPF trade-off over confidence thresholding. The annotated comparisons indicate 1.3× TPF at matched accuracy or 10.6 percentage points of accuracy at matched TPF.
The 3B and 14B instruct rows report linear-SS TPF of 4.36 and 5.96 with mean accuracies of 55.00 and 66.36. The 3B comparison includes Qwen3-4B, so it is not a parameter-matched 3B comparison.
The proposed 8B base model obtains 72.36 mean accuracy and 4.67 TPF in linear SS, compared with Qwen3-8B base's 71.58 and 1.00. Its four mode rows show that accuracy and TPF do not change identically across decoding choices.
6.4. Benchmark VLMs
Across eight VLM benchmark columns, AR and linear-SS modes average 59.5 and 59.4 accuracy, while diffusion averages 57.9 with denoising threshold τ = 0.9. Linear SS reaches 3.63 TPF over all samples and 7.45 TPF for responses longer than 200 tokens; the latter is a length-filtered TPF figure, not all-sample throughput.
VLM linear SS averages 59.4 accuracy against AR's 59.5. Its 7.45 TPF applies only to responses exceeding 200 tokens, whereas the all-sample linear-SS figure is 3.63 TPF.
6.5. Inference Efficiency
SGLang tests compare the 8B model with Qwen3-8B-Eagle3 on SPEED-Bench math, coding, reasoning, and multilingual prompts, with generation capped at 1024 tokens. At concurrency 1 on GB200, linear SS produces 851 tok/sec versus 256 for AR and 354 for Eagle3; an optimized kernel reaches 1015 tok/sec, while 1471 tok/sec for SOL is a projection.
- At draft length 31 across 11 categories, LoRA self-speculation averages 6.82 accepted tokens, versus 2.75 for Eagle3 and 4.24 for Qwen3-9B-MTP; the four-category LoRA average is 8.69. On DGX Spark, linear SS reaches 77.5 tok/sec in FP8 (3.14× AR) and 112.5 tok/sec in INT4 (2.69× AR). Reported concurrency curves show a different system-versus-per-user throughput trade-off at higher load.
The device panels distinguish measured concurrency-1 throughput from hatched SOL estimates. On GB200, measured linear SS reaches 851 tok/sec, or 1015 with an optimized kernel, versus 256 for AR; the other panels separately report RTX Pro 6000 and DGX Spark results.
At draft length 31, LoRA self-speculation averages 6.82 accepted tokens across 11 categories versus 2.75 for Eagle3 and 4.24 for MTP. Its 8.69 four-category average reflects the selected coding, math, reasoning, and multilingual subset rather than the full benchmark.
8. Insights and Future Directions
Joint training yields three inference modes without mode-specific architectural changes; the coefficient sweep and matched-token ablation show preserved mean AR accuracy, while the LoRA ablation and device tests support linear self-speculation’s practical low-concurrency advantage over the evaluated Eagle3 and MTP comparators. Global loss averaging and AR initialization are part of the reported training recipe, but the cumulative ablation does not isolate their effects under identical total training budgets.
- Closing the oracle-to-sampler gap, improving drafter–verifier alignment, verifying beyond a contiguous AR prefix, and seeking segment-level parallelism are proposed directions. The SOL procedure requires costly simulations against the model’s own serial output, so its 76.5% TPF difference does not establish a deployed speed gain.
Appendix
- Appendix A describes a 4-layer, 384-hidden-dimension sampler with 4.8M parameters and 144-dimensional per-position features, trained using approximately 20M denoising trajectories; an MLP-only ablation loses 10 percentage points of accuracy–TPF AUC. Appendix B illustrates diffusion-drafter LoRA training. Appendix C specifies a quadratic layout with k² inserted masks and an optional verifier interpolating AR and diffusion predictive distributions.
Brief Thoughts
The controlled AR-loss comparison and the measured low-concurrency throughput results support different parts of the paper’s case for a shared AR–diffusion model. The SOL construction usefully separates potential acceptance length from forward-pass cost, but its oracle search, distinct correctness target, and the accuracy loss at larger blocks limit what it establishes about a practical diffusion sampler.