LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model | Summary
22 Apr 2026 | Paper Review Diffusion Language Models Unified Multimodal Models Semantic TokenizationContents
- Summary
- 1 Introduction
- 2 Model Design
- 3 Data Preparation
- 4 Model Training
- 5 Experiments
- 6 Conclusion and Future Directions
- Appendix
- Brief Thoughts
This article explains the key points of LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model.
- 2026-04-22 (arXiv)
- AI, Inclusion, Bie, Tiwei, Chen, Haoxing, Chen, Tieyuan, Cheng, Zhenglin, Cui, Long, Gan, Kai, Huang, Zhicheng, Lan, Zhenzhong, Li, Haoquan, et al.
- AGI Research Center, Inclusion AI
- Paper
- Github
Summary
- LLaDA2.0-Uni uses shared discrete semantic image tokens for visual understanding, image generation, editing, and interleaved text–image outputs. Its architecture combines SigLIP-VQ, a 16B Mixture-of-Experts discrete diffusion language model, and a separately trained 6B diffusion decoder that reconstructs images from semantic tokens.
- The model achieves 64.1 on MMStar, 0.89 on GenEval, 79.63 on UniGenBench, and 47.1 on MICo-Bench, although understanding and editing results do not uniformly match the strongest listed specialists. SPRINT raises average backbone throughput from 24.3 to 39.8 TPS with a reported average score decrease of 0.6; decoder distillation reduces decoder latency from 32.95 to 2.90 s/img under the paper’s single-GPU measurement conditions.
- Fine visual detail and dense-text rendering remain limitations, and interleaved reasoning is supported only by selected examples rather than quantitative evaluation. Interleaved generation is evaluated on the authors’ 150-sample InterGen benchmark, where gains over Emu3.5 are modest and explanation results depend on the judge. Figure 1 summarizes benchmark results, while Figures 2 and 3 illustrate generation, editing, and interleaved outputs.
The radial comparisons summarize understanding on the left and generation/editing on the right. Exact values do not always match the detailed tables: Lumina-DiMOO’s GenEval is displayed as 0.89 versus 0.88 in Table 3, and LLaDA2.0-Uni’s MathVision is displayed as 28.0 versus 26.7 in Table 2. LLaDA-o’s displayed GenEval value of 0.86 agrees with Table 3.
The gallery includes photographic portraits, illustrations, objects, and English and Chinese text embedded in images. It illustrates stylistic range and visible text rendering, but provides no prompts, sampling protocol, or failure-rate information establishing representative generation quality.
Before-and-after pairs illustrate localized changes, while the multi-reference example combines clothing, accessories, a person, and a scene. The lower panels demonstrate alternating descriptions and images, including chess-state visualizations; these are selected capability examples rather than quantitative evaluations.
1 Introduction
The paper addresses limitations of unified masked-diffusion models relative to autoregressive and hybrid multimodal systems. It attributes prior weaknesses to reconstruction-oriented visual tokenizers, excessive image compression, unconstrained bidirectional attention, and fixed-length understanding outputs. LLaDA2.0-Uni instead shares semantic image tokens across understanding and generation, predicts text and images through block-wise masked diffusion, and delegates image reconstruction to a dedicated decoder.
- The proposed contributions are the shared semantic-token architecture, interleaved generation and reasoning, inference acceleration, and broad benchmark evaluation.
- Resources listed in the PDF are https://github.com/inclusionAI/LLaDA2.0-Uni and https://huggingface.co/inclusionAI/LLaDA2.0-Uni.
2 Model Design
The model design separates semantic sequence modeling from image reconstruction. Text and visual tokens share the dLLM backbone and its mask-prediction objective, while the diffusion decoder converts generated semantic visual tokens into images. Unification therefore concerns the backbone representation and objective, not a single loss jointly training every component.
2.1 Motivation and Design Principles
The central design principle is to use fully discrete semantic image tokens for both visual understanding and generation. The authors contrast this with reconstruction-based VQ tokens and systems that use separate understanding and generation encoders. Sharing tokens avoids separate visual representations for the two tasks, although reconstruction still requires a specialized image decoder.
2.2 Architecture
The architecture consists of a semantic discrete tokenizer, a modality-agnostic MoE diffusion language model, and a diffusion decoder. Figure 4 shows understanding, editing, and text-to-image generation using the same backbone, with task-specific masking and separate text or image decoding.
- 2.2.1 Semantic Discrete Tokenizer: SigLIP-VQ builds on X-Omni and uses a pre-trained SigLIP2-g ViT with dynamic-resolution support. It is trained on understanding tasks rather than pixel reconstruction and has a codebook vocabulary of 16,384 with dimensionality 2,048.
- 2.2.2 Diffusion Large Language Model: LLaDA-2.0-mini supplies 16B total MoE parameters. Visual codebook entries and special tokens extend its vocabulary; existing language embeddings and prediction-head weights are retained, while new visual embeddings are randomly initialized.
- Block-wise attention permits bidirectional processing within blocks while conditioning on preceding blocks. The authors motivate this partly through the autoregressive bias inherited by SigLIP-VQ tokens from alignment with Qwen2.5. Original 1D RoPE is retained, with <height> and <width> tokens, such as <imgsize 512>, representing image dimensions before flattened visual sequences.
- 2.2.3 Diffusion Decoder: a 6B Z-Image-Base-derived model receives semantic visual tokens as its sole conditioning signal, without an additional text prompt. It performs 2× super-resolution using upsampled semantic tokens and is distilled from 50-step sampling with CFG to 8-step CFG-free inference.
Text and SigLIP-VQ image tokens enter the same LLaDA2.0-mini backbone. Understanding produces text tokens, while editing and generation produce image tokens passed to the diffusion decoder; height and width tokens encode image dimensions.
2.3 Training-free Inference Acceleration
SPRINT, Sparse Prefix Retention with Inference-time Non-uniform Token Unmasking, accelerates block-wise diffusion without retraining. It prunes the prefix KV cache used in subsequent denoising steps and accepts confident predictions adaptively to reduce the number of steps. Standard decoding requires (B \times T) forward passes for B blocks with T denoising steps each.
- Sparse Prefix Retention performs a full first step for each block, then ranks prefix positions using (s_i = \alpha\bar I_i + (1-\alpha)c_i) with α = 0.5. Ī_i is the mean-normalized key norm, and c_i is top-1 softmax confidence; separate text and image keep ratios are intended to preserve instruction and reasoning tokens while pruning redundant image tokens.
- The stated settings are rtext = 1.0, rimg = 0.8 and rtext = rimg = 1.0. The text also states a global keep ratio r = 0.5 for both settings without explaining its relationship to the modality-specific ratios.
- Non-uniform Token Unmasking accepts every masked position whose confidence exceeds τ, examining τ ∈ {0.93, 0.95}. A minimum of ⌈m/(T − t)⌉ acceptances per step guarantees termination rather than relying solely on the threshold.
3 Data Preparation
Data preparation covers multimodal understanding, image generation, image editing, interleaved sequences, and reasoning-augmented examples. The pipeline combines public datasets, web images, synthesized editing pairs, video-derived sequences, and model-assisted filtering and captioning.
3.1 Multimodal Understanding
Understanding pre-training combines image-caption data with OCR, grounding, counting, world knowledge, reasoning, and text-only material. The SFT collection contains approximately 60 million samples with a 1:5 ratio of text-only to multimodal data, covering dialogue, multiple images, VQA, charts, tables, and mathematical reasoning.
- OCR supervision uses PaddleOCR pseudo-labels refined by Qwen3-VL. Grounding data from Objects365 and RefCOCO undergoes confidence filtering and Qwen3-VL-235B-A22B verification; coordinates are normalized to [0, 1000], and verified bounding boxes also provide counting examples.
- Text-only data comes from Ling2.0 and LLaDA2.0. For SFT quality control, Qwen3-VL removes low-information queries and rewrites ambiguous ones, while rules correct response artifacts and GPT-OSS filters semantic biases against ground-truth references.
3.2 Image Generation
The generation corpus begins with over 200 million web images and retains 140 million after metadata, aesthetic, and quality filtering. The collection increases the proportion of human-body and rendered-text images to address difficult generation cases. Qwen3-VL-235B-22B recaptions the images, incorporating informative original web descriptions to retain specific real-world knowledge.
- Metadata filtering removes images with a shortest side below 512 pixels and images satisfying (Height×Width)/Filesize < 0.15; the PDF does not specify the filesize unit.
- Aesthetic filtering rejects ArtiMuse scores below 60, and quality filtering rejects DeQA-Score values under 4.0.
3.3 Image Editing
Editing data combines X2Edit, OmniEdit, Nano-consistent-150k, Pico-Banana, UniWorld, StructVisuals, UnicEdit, and CrispEdit with synthesized pairs derived from the generation corpus. Qwen3-VL-235B-22B removes edits with no observable change or visible artifacts and rewrites inaccurate or vague instructions to describe the actual transformation.
3.4 Interleaved Data
Interleaved training data is constructed from Koala36M videos and yields 6M refined clips. Frames are sampled every 5 seconds to form sequences of 2–6 images; Qwen3-VL-235B-A22B describes actions and scene changes and generates corresponding user instructions for SFT.
- Clips longer than 30 seconds or shorter than 10 seconds are discarded. Quality filtering retains the top 50% by aesthetic score (> 4.0) and clarity (> 0.7), while motion scores must exceed 4 to discourage static-image outputs.
- The paper reports that filtering removes approximately 75% of raw data but does not state the starting subset size needed to reconcile that percentage with 6M retained clips from Koala36M.
3.5 Reasoning-Augmented Data
Approximately 8M SFT samples from Flux-6M, Zebra-CoT, and Weave provide reasoning-based image generation and interleaved reasoning supervision. These examples support textual chain-of-thought before image synthesis and multi-step reasoning across alternating text and images. The paper does not separately ablate this data mixture.
4 Model Training
Backbone training progresses through vision–language alignment, multi-task pre-training, and supervised fine-tuning. Optimization combines block diffusion, MoE load balancing, and SFT-specific masking and loss reweighting. The image decoder follows a separate flow-matching and distillation pipeline.
4.1 Training Recipe
Table 1 specifies 100B training tokens for S0, 210B for S1, and 80B for S2. Context length is 8192 in S0 and S1 and expands from 8192 to 16384 during S2. Generation resolution progresses from 256 to 512 in S0, remains 512 in later backbone stages, and reaches 1024 through the diffusion decoder.
- Stage 0: Vision–Language Alignment uses image-caption pairs, visual knowledge, and pure text. Generation tasks mask only image tokens, while understanding tasks mask only text tokens; generation moves from 256 × 256 (∼256 tokens) to 512 × 512 (∼1024 tokens).
- Understanding uses arbitrary-resolution processing with a maximum edge of 800; the prose describes 800 × 800 inputs as approximately 2048 tokens.
- Stage 1: Multi-task Pre-training introduces OCR, counting, grounding, video, editing, reference-based style transfer, subject-driven generation, controllable generation, and multi-view tasks.
- Stage 2: Supervised Fine-tuning first develops instruction following at 8k context, then extends to 16k for complex visual reasoning and generation, including interleaved and CoT-augmented tasks.
The stage comparison documents the training curriculum: 100B, 210B, and 80B tokens, with context extended to 16384 in S2. It distinguishes 512-resolution backbone generation from 1024-resolution decoder output and shows the introduction of editing, interleaved generation, and reasoning data.
4.2 Pre-Training Optimization
Pre-training uses the Block Diffusion Language Model objective: masked tokens in each current block are predicted from preceding clean blocks and the noisy current block. The loss includes a masked-position indicator and the diffusion-derived time weight (-\frac{\alpha’_t}{1-\alpha_t}). This combines parallel prediction within blocks with ordered conditioning across blocks.
- Equation (3) defines K = Ltotal/LB blocks, where LB is the block size. Only masked positions contribute prediction terms.
- MoE optimization uses auxiliary-loss-free load balancing and scales routing gate outputs by 2.5 for numerical stability. Equation (4) applies an RMSNorm-style normalization to expert-bias updates using the current expert-load distribution and a uniform target distribution.
4.3 Supervised Fine-Tuning Optimization
SFT conditions the block-diffusion objective on the input prompt and retains the same MoE balancing strategy. Because sample lengths vary by up to two orders of magnitude, Mask Token Reweighting Loss uses an inverse-square-root masked-token count to balance the influence of long and short samples. Complementary masking provides paired corrupted sequences with inverse masks.
- Mask Token Reweighting Loss: Equation (6) defines LMTRS = (∑j βj LSFT^(j))/(∑j βj), with βj equal to the inverse square root of the number of masked tokens in sample j. The stated motivation is to avoid long-sequence dominance under token averaging and brevity incentives under sample averaging.
- Complementary Masking: each clean sequence produces a primary noisy instance and an inverse-mask counterpart, so every position is uncorrupted exactly once per pair. The strategy is intended to improve information utilization and reduce token-level sampling bias.
4.4 Diffusion Decoder Training
The decoder is trained with flow matching, minimizing squared error between predicted and target velocity while conditioning on semantic visual tokens. Warm-up freezes the semantic processor; multi-domain generalization then unfreezes all parameters, followed by high-quality-data refinement. Consistency-based distillation enables 8-step CFG-free decoding.
- Equation (7) gives LFM(θ) = E[‖vθ,t(xt, z) − vt‖²₂], where z is the semantic-token conditioning signal.
- Few-step Generation: distillation adds an auxiliary final projection layer that is discarded at inference. Equation (8) combines flow matching with a consistency term involving a stop-gradient output and its time derivative.
- The time derivative is a Jacobian-vector product approximated with the second-order difference technique from UCGM.
4.5 Infrastructure for Training Efficiency
Training efficiency relies on offline visual-token extraction, offline sequence packing, and distributed execution through dFactory built on VeOmni. Images are processed through the frozen tokenizer before backbone training, and their token indices are stored on disk. Figure 5 illustrates packing short samples into fixed-length sequences to reduce padding across tasks with heterogeneous lengths.
- These choices avoid repeated visual-encoder computation and increase the fraction of useful tokens in each batch. The section does not report controlled throughput comparisons for the individual infrastructure changes.
Samples from different tasks occupy portions of the same fixed-length packed sequence rather than being padded individually to the longest sample. This explains the intended utilization benefit, although the paper does not quantify packing’s isolated speedup.
5 Experiments
Experiments cover multimodal understanding, general and text-centric image generation, reasoning-informed generation, instruction-based and multi-reference editing, and interleaved outputs. Comparisons include specialists, autoregressive unified models, hybrid autoregressive–diffusion systems, and discrete-diffusion models. Separate acceleration analyses measure SPRINT and decoder distillation.
5.1 Multimodal Understanding
Table 2 shows competitive understanding performance, but not uniform parity with specialized VLMs or superiority over other unified diffusion models. LLaDA2.0-Uni scores 64.1 on MMStar and 86.0 on CountBench, above Qwen2.5-VL-7B’s 63.9 and 84.9, while trailing it on OCRBench (75.7 versus 84.2), DocVQA (89.5 versus 94.9), and V* (61.8 versus 80.1).
- 5.1.1 Evaluation Settings: the paper lists 21 named benchmarks across general VQA, reasoning, OCR/documents, and other tasks, comparing with Qwen2.5-VL-7B, LLaDA-V, BAGEL, InternVL-U, Lumina-DiMOO, and LLaDA-o. Table 2 has 21 result rows covering 20 named benchmarks: MMBench is split into English and Chinese, while the listed MathVerse has no displayed result.
- 5.1.2 Multimodal Understanding Performance: MathVista is 68.1 versus Qwen2.5-VL-7B’s 68.2, and MathVision is 26.7 versus 22.4. MMMU is 50.1, below Lumina-DiMOO’s 58.6, and ChartQA is 80.1, below LLaDA-o’s 87.9.
- The results show broad understanding capability and substantially stronger OCR/document performance than Lumina-DiMOO. They do not support the paper’s blanket claim of improvement over both unified diffusion baselines across all major categories.
The results support competitive general understanding and much stronger OCR performance than Lumina-DiMOO, but not universal specialist parity. LLaDA2.0-Uni leads the displayed MathVision comparison at 26.7 yet trails Qwen2.5-VL-7B on OCRBench, DocVQA, and V*. The table has 21 rows, with MMBench split by language and no MathVerse result.
5.2 Text-to-Image Generation
LLaDA2.0-Uni achieves 0.89 on GenEval, 87.76 on DPG, 0.505 on OneIG-EN, and 79.63 on UniGenBench. These are the highest overall scores among the unified models listed in the respective tables, but specialists remain stronger on several text-rendering and overall-generation metrics.
- 5.2.1 Evaluation Settings: GenEval, DPG-Bench, One-IG Bench, and UniGenBench assess general generation; CVTG-2K assesses text rendering; WISE-Bench assesses reasoning-informed generation. Tables 3–8 compare generation-only specialists with autoregressive, hybrid, and diffusion-based unified models.
- 5.2.2 General Image Generation Performance: Table 3 reports Position 0.90 and Attribute Binding 0.84, but Counting is 0.73 versus Seedream 3.0’s 0.91. Contrary to the prose, Position is not the highest listed score: LLaDA-o reaches 0.69 and Lumina-DiMOO 0.85, so the claim is supported by Table 3. Table 4 reports Entity 93.55 and Other 93.18; the prose instead gives Other as 94.04.
- On OneIG-EN, Alignment is 0.882 and Reasoning is 0.323, but Text is 0.661 versus Qwen-Image’s 0.891 and Z-Image-Turbo’s 0.994. On UniGenBench, Logic is 63.99 and Layout is 90.30, while Text is 31.23 versus Qwen-Image’s 72.13.
- 5.2.3 Text-Centric Image Generation Performance: CVTG-2K average word accuracy is 0.765, declining from 0.788 for 2 regions to 0.746 for 5 regions. The decline is smaller than for the listed unified baselines, but average accuracy remains below LongCat-Image’s 0.865.
- 5.2.4 Reasoning-Informed Image Generation Performance: WISE rises from 0.68 to 0.78 with thinking. This is an absolute increase of 0.10 score points; the paper’s description of an additional 10% improvement does not distinguish absolute from relative change.
LLaDA2.0-Uni has the highest displayed GenEval overall score, 0.89, and the highest Position score, 0.90, with Attribute Binding at 0.84. Counting remains 0.73, indicating that strong spatial and attribute composition does not imply equally strong numerosity control.
The 87.76 overall score leads the listed unified models but remains below Qwen-Image’s 88.32. The table reports Other as 93.18 rather than the prose’s 94.04 and shows the highest displayed Entity score at 93.55.
Alignment 0.882 matches Qwen-Image, and Reasoning 0.323 exceeds the listed baselines. Text 0.661 is lower than several specialists and some unified competitors, leaving overall performance at 0.505 below the strongest generation-only results.
The 79.63 overall result exceeds all listed models, with strong Logic 63.99, Composition 85.82, and Layout 90.30 scores. Text 31.23 remains substantially below leading specialists, so the overall result masks an uneven capability profile.
Word accuracy falls from 0.788 to 0.746 as text regions increase from two to five, a smaller decline than for the listed unified baselines. Its average 0.765 remains below LongCat-Image’s 0.865 and Z-Image-Turbo’s 0.858.
Adding thinking raises the overall WISE score from 0.68 to 0.78, with particularly large changes in Cultural knowledge and Chemistry. The table supports better scores under reasoning mode but does not report its additional inference cost.
5.3 Image Editing
Editing results are strongest for multi-reference composition, while general editing remains below leading specialists. LLaDA2.0-Uni reaches 3.92 on ImgEdit, the best score among the listed unified models, and 47.1 on MICo-Bench, above every listed baseline. GEdit overall scores are 6.61 in English and 6.66 in Chinese, below the strongest specialized editing models.
- 5.3.1 Evaluation Settings: ImgEdit-Bench and GEdit-Bench evaluate instruction-based editing, while MICo-Bench evaluates multi-reference composition. Baselines include FLUX.1 Kontext, Step1X-Edit, Qwen-Image-Edit, Z-Image-Edit, BAGEL, OmniGen2, InternVL-U, and Lumina-DiMOO.
- 5.3.2 General Image Editing Performance: ImgEdit Adjust and Hybrid scores are 4.16 and 3.97. On GEdit, perceptual quality is 7.52 in English and 7.67 in Chinese, but semantic consistency is 6.68 and 6.63; the English overall score also trails InternVL-U’s 6.66.
- 5.3.3 Multi-Reference Image Editing Performance: MICo-Bench results are 51.0 for Object, 32.8 for Person, 46.0 for HOI, and 54.4 for De&Re. Object performance is below Qwen-Image-Edit’s 52.4, but the remaining categories and overall score lead the listed comparisons.
- The MICo-Bench result supports strong composition capability, but the experiment does not isolate architecture from training-data effects.
The overall score of 3.92 leads the unified group, and Hybrid 3.97 exceeds every listed model. General editing still trails Qwen-Image-Edit’s 4.35 overall, while Extract is relatively weak at 2.40.
LLaDA2.0-Uni’s perceptual-quality scores, 7.52 in English and 7.67 in Chinese, exceed BAGEL’s, but its semantic-consistency scores are lower. The results distinguish visually plausible edits from faithful execution of editing instructions.
The 47.1 overall result exceeds Qwen-Image-Edit’s 35.9, BAGEL’s 34.4, and OmniGen2’s 33.8. LLaDA2.0-Uni leads Person, HOI, and De&Re, while Qwen-Image-Edit retains a small Object-category advantage.
5.4 Interleaved
The authors introduce InterGen, a 150-sample benchmark for Story Telling, Explanation, and Event Forecasting, judged by Gemini-3 and Qwen3-VL. LLaDA2.0-Uni modestly outperforms Emu3.5 on storytelling and forecasting under both judges, while explanation results depend on the judge. Figures 6–8 present the benchmark design, interleaved generation examples, and preliminary reasoning demonstrations.
- 5.4.1 Interleaved Generation: judges assess text coherence, text–image alignment, and ID consistency. Table 12 reports Story Telling scores of 6.42/7.02 versus Emu3.5’s 6.28/6.83, and Event Forecasting scores of 5.19/5.94 versus 5.08/5.75, for Gemini/Qwen3-VL respectively.
- Explanation scores are 6.22/6.35 versus 6.19/6.48, so improvement is not consistent across judges. The comparison includes only Emu3.5; the authors state that other relevant systems were not yet open-source.
- 5.4.2 Interleaved Reasoning: Figure 8 alternates generated physics diagrams or chess positions with textual reasoning. These examples demonstrate the output format but do not establish aggregate accuracy, reasoning faithfulness, or an advantage from generating intermediate images.
- The small author-constructed benchmark and absence of reported uncertainty estimates limit conclusions about the modest InterGen score differences.
InterGen organizes tasks into Story Telling, Explanation, and Event Forecasting, including cooking procedures and action anticipation. The evaluator panel identifies coherence, alignment, and identity consistency as VLM-judged dimensions rather than direct ground-truth matches.
A character sequence illustrates recurring appearance and scene continuity, while a cooking sequence pairs procedural text with images. The character prompt requests 6 images, but only 4 are displayed, so the panel does not establish complete compliance with that request or reliability over longer sequences.
Both judges score LLaDA2.0-Uni above Emu3.5 for storytelling and event forecasting by modest margins. Explanation is mixed: Gemini scores it slightly higher, while Qwen3-VL scores it lower, precluding a claim of uniform superiority.
The physics example generates force diagrams alongside a textual solution, while the chess example generates candidate board states. The displayed physics response reports acceleration of approximately 5.98 m/s² and tension of approximately 23.67 N; these are individual model outputs, not aggregate evidence of reasoning accuracy.
5.5 Ablation Study
Acceleration experiments show separate trade-offs at the backbone and image-decoder levels. SPRINT increases average throughput by a reported 1.6× while reducing the reported average benchmark score from 76.3 to 75.7. Decoder Turbo provides a reported 11.4× decoder-only speedup with small score changes across five generation benchmarks.
- 5.5.1 Analysis of SPRINT Acceleration: Table 13 measures nine benchmarks without SGLang. Average throughput rises from 24.3 to 39.8 TPS; OCRBench falls from 75.7 to 73.4 and DPG from 87.76 to 86.27, while MMMU rises from 50.1 to 52.5.
- DocVQA throughput rises from 8.0 to 27.6 TPS, a reported 3.5× gain. The authors suggest confidence-adaptive acceptance reallocates refinement to difficult positions, but the experiment does not separately ablate prefix pruning and adaptive unmasking. SGLang integration is described as ongoing work.
- 5.5.2 Analysis of Diffusion Decoder: Table 14 measures a single GPU at 1024 × 1024 resolution, batch size 1, and BF16 precision; the GPU model is not specified. Decoding drops from 32.95 to 2.90 s/img when moving from 50 to 8 steps.
- Turbo scores are GenEval 0.87 versus 0.89, DPG 87.24 versus 87.76, UniGenBench 79.76 versus 79.63, OneIG-EN 0.500 versus 0.505, and WISE 0.68 for both. Figure 9 shows similar selected outputs, not a controlled perceptual-equivalence test.
- The decoder speedup excludes backbone token generation and should not be interpreted as an 11.4× end-to-end speedup. The SPRINT average combines benchmarks with different score meanings rather than measuring quality on a common calibrated scale.
SPRINT raises reported average throughput from 24.3 to 39.8 TPS while the reported average score falls by 0.6. Task-level results expose the trade-off: OCRBench loses 2.3 points, MMMU gains 2.4, and DocVQA receives the largest reported throughput multiplier. GenEval is expressed on a 100-point scale in this table.
Paired outputs from the 50-step and 8-step decoders retain similar subjects, layouts, and styles across four examples. Visible local differences mean the figure supports qualitative preservation rather than perceptual indistinguishability.
Decoder latency decreases from 32.95 to 2.90 s/img under single-GPU, 1024 × 1024, batch-size-1, BF16 conditions. GenEval and DPG decline slightly, UniGenBench increases slightly, and WISE is unchanged. These timings exclude the backbone token-generation pipeline.
6 Conclusion and Future Directions
The conclusion identifies semantic discrete tokens and block-wise diffusion as the basis for unified understanding and generation. The authors acknowledge that SigLIP-VQ loses fine-grained visual detail, particularly relevant to editing, and propose scaling interleaved data and model capacity. Reinforcement learning is described as ongoing work with unresolved optimization challenges, not as an evaluated contribution.
Appendix
- The supplied 25-page PDF contains no appendix. Pages 20–25 are References, with no additional appendix experiments, implementation details, or supplementary results.
Brief Thoughts
The main contribution is the division between shared semantic discrete modeling and continuous image reconstruction. The benchmark results support the usefulness of this design, but controlled tokenizer, attention, and data-mixture ablations are absent, so the gains cannot be attributed to a particular component. The decoder comparison provides the clearest efficiency evidence under stated resolution, batch-size, and precision conditions, although hardware identity is missing. SPRINT improves throughput with measurable OCR and generation losses, and its pruning configuration is incompletely explained. Interleaved reasoning remains exploratory: selected chess and physics examples do not establish reliable visual reasoning or an advantage over text-only reasoning. Numerical inconsistencies include DPG Other in the prose versus Table 4 and Figure 1’s GenEval and MathVision displays versus the detailed tables. The evaluation settings also list MathVerse without a result. Quantitative judgments here use the detailed tables.