Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling | Summary
02 Jul 2026 | Paper Review Diffusion Acceleration Flow Matching Text-to-Image Generation Super-ResolutionContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Method
- 4 Empirical Results
- 5 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling.
- 2026-07-02 (arXiv)
- Zheng, Xingyu, Liu, Xianglong, Ding, Yifu, Feng, Weilun, Lin, Junqing, Guo, Jinyang, Qin, Haotong.
- State Key Laboratory of Complex & Critical Software Environment, Beihang University, School of Computer Science and Engineering, Beihang University, Nanyang Technological University, State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, University of Chinese Academy of Sciences, University of Science and Technology of China, School of Artificial Intelligence, Beihang University, ETH Zürich
- Paper
- Github
Summary
- MrFlow accelerates inference with pretrained text-to-image flow-matching models through low-resolution structure generation, pixel-space super-resolution, low-strength latent noising, and brief high-resolution refinement. The sampling pipeline needs no additional training or runtime region identification, but uses a pretrained Real-ESRGAN model.
- At 1024 × 1024, Qwen-Image MrFlow with 12 low-resolution and 1 high-resolution steps reaches 10.3× end-to-end speedup; OneIG-En and OneIG-Zh change from 0.53/0.53 to 0.52/0.51. The corresponding FLUX.1-dev setting reaches 8.25× speedup.
- Combining the pipeline with pretrained Pi-Flow weights reaches 25.1× speedup on Qwen-Image. That combination requires a separately trained distilled model, unlike MrFlow applied to the original flow model.
The Qwen-Image examples juxtapose native output with several accelerators, including MrFlow at 10.3×. Dashed divisions distinguish the training-free comparisons from Pi-Flow and the 25.1× MrFlow combination that uses distilled weights; the images permit inspection of local output differences beyond benchmark scores.
1 Introduction
The paper reports up to 47s of diffusion sampling for Qwen-Image-20B at 1024 × 1024 on an Nvidia A100. Timestep distillation reduces the number of model evaluations but requires training; training-free caching and token reduction offer smaller speedups in the comparisons discussed here. Earlier multi-resolution accelerators upsample in latent or frequency space, where the authors report blur or artifacts in some outputs.
MrFlow assigns prompt-conditioned layout to low-resolution sampling and local correction to high-resolution sampling. Figure 2 contrasts its pixel-space super-resolution and VAE re-encoding with approaches that enlarge the latent representation directly.
MrFlow’s path runs from low-resolution latent sampling through VAE decoding, pixel-space GAN super-resolution, VAE re-encoding, small-noise injection, and high-resolution denoising. The separate native and latent-upsampling paths identify where its representation change occurs.
2 Related Work
Related work covers flow matching, timestep reduction, feature caching, token compression, multi-resolution generation, super-resolution, and image repainting. MrFlow targets acceleration at the model’s evaluation resolution: it enlarges a decoded image using pretrained pixel-space super-resolution, then re-encodes, noises, and refines the latent rather than directly enlarging it in latent or frequency space.
3 Method
The pipeline samples a low-resolution latent, VAE-decodes it, applies pixel-space super-resolution, VAE-encodes the enlarged image, adds noise, performs high-resolution flow sampling, and decodes the result. A typical setting uses 12 low-resolution Euler steps, noise strength 0.1, and 1 high-resolution step; both sampling stages use the same pretrained velocity network and text condition.
3.1 Low-Resolution Structure Generation
Low-resolution generation starts from Gaussian latent noise and uses Euler integration of the pretrained flow ODE to produce a latent for VAE decoding. Halving each image side reduces the image-token count by four and gives about 4× measured per-step acceleration; the paper also tests whether fewer steps can establish the global structure.
The bookstore example displays each super-resolution output above its high-resolution-refined result. Enlarged sign details expose blur or lettering differences that the close aggregate super-resolution scores do not resolve.
3.2 Pixel-Space Super-Resolution
Pretrained Real-ESRGAN enlarges the decoded image in pixel space, typically by ×2. This choice supplies sharp details while retaining the coarse layout needed for refinement. Because the super-resolution network is not conditioned on the prompt, the enlarged image can still contain artifacts or shifted character strokes.
3.3 Low-Strength Noise for High-Frequency Resampling
The enlarged image is VAE-encoded and mixed with Gaussian noise at the refinement timestep, typically with σt in [0.1, 0.15]. Under a local Gaussian prior and high-frequency-error assumption, σt ≥ √λhf/(1 + √λhf) makes injected noise power at least as large as clean-signal power in the designated high-frequency directions. The bound establishes a resampling scale, not a guarantee that short refinement corrects arbitrary super-resolution errors or diffuse blur.
The watchmaker example holds the 20-step low-resolution stage fixed and varies high-resolution refinement from zero to two steps. Notebook lettering improves most between zero and one refinement step; the second produces a smaller visible difference.
3.4 High-Resolution Detail Refinement
High-resolution sampling begins from the noised super-resolution latent and defaults to one Euler step before final VAE decoding. For Qwen-Image at strength 0.1, its one-step output has CLIP image similarity 0.9974 to an 8-step refinement reference, versus 0.9999 for five steps. The paper relates this result to measured velocity changes near the clean endpoint.
4 Empirical Results
The main experiments assess generation metrics and end-to-end inference speed on FLUX.1-dev and Qwen-Image-20B. They separate training-free methods from comparisons using released timestep-distillation or latent-upsampling weights.
4.1 Implementation Details
Quantitative evaluation uses 1024 × 1024 images and Geneval, DPG-Bench, and OneIG-Bench. Each experiment runs on a single Nvidia A100; end-to-end timing includes text encoding, noise generation, VAE operations, super-resolution, and diffusion forward passes. Native Qwen-Image uses two forward passes per step for true CFG, while methods needing pretrained weights use released weights without retraining in these experiments.
At the aggressive Qwen-Image operating point, MrFlow records 10.3× speedup and Geneval 0.86; the two feature-cache rows near 9× record 0.26 and 0.09. Its FLUX.1-dev (12, 1) row records 8.25× speedup, Geneval 0.63, and DPG 81.65.
4.2 Main Results
In Table 1, Qwen-Image MrFlow at (12, 1)×2 reaches 10.3× speedup with Geneval 0.86, DPG 87.10, OneIG-En 0.52, and OneIG-Zh 0.51, against native scores of 0.88, 88.67, 0.53, and 0.53. At nearby aggressive speedups, DB-Taylor records Geneval 0.09 at 9.34× and RALU records 0.84 at 9.51×. On FLUX.1-dev, MrFlow reaches 8.25× with Geneval 0.63 versus 0.66 natively.
The Qwen-Image rows distinguish training-free MrFlow from MrFlow† using pretrained Pi-Flow weights. Relative to Pi-Flow, MrFlow† raises speedup from 19.6× to 25.1× while retaining Geneval 0.85 and recording slightly lower OneIG scores.
With pretrained Pi-Flow weights, Qwen-Image MrFlow† reaches 25.1× speedup versus 19.6× for Pi-Flow alone. Both record Geneval 0.85; MrFlow† records OneIG-En/OneIG-Zh 0.52/0.51 versus Pi-Flow’s 0.53/0.52. The combination adds no training to the pipeline but depends on previously distilled weights.
The plots place FLUX.1-dev and Qwen-Image Geneval scores against measured speedup for the tested methods and step settings. The closely grouped MrFlow curves with different high-resolution step counts are consistent with the reported ablation that additional refinement steps contribute little to these metric trade-offs.
4.3 Ablation Study
In the tested step-count sweep, more low-resolution steps improve the Geneval–speedup trade-off, while increasing high-resolution steps shows little metric benefit. The visual comparison shows clearer notebook lettering after one refinement step than with none and a smaller visible difference after a second.
Interpolation and Real-ESRGAN both record OneIG-En 0.52 and OneIG-Zh 0.51 at (12, 1)×2, with speedups of 10.7× and 10.3×; SwinIR records 4.67×. The near-equal OneIG values do not by themselves distinguish the local visual defects shown in Figure 3.
The super-resolution ablation compares interpolation, SwinIR, OSEDiff, and Real-ESRGAN. At (12, 1)×2, interpolation and Real-ESRGAN both record OneIG-En 0.52 and OneIG-Zh 0.51, but Figure 3 shows blur with interpolation and SwinIR and character inaccuracies with OSEDiff. The authors select Real-ESRGAN for the tested balance of visual detail and speed.
5 Conclusion
MrFlow moves most sampling to cheaper low resolution while keeping a short high-resolution correction stage. The reported Qwen-Image training-free operating point reaches 10.3× end-to-end speedup; pairing the pipeline with pretrained timestep distillation reaches 25.1×, a distinct setting with training-dependent weights.
Appendix
- Appendix A expands the comparisons with acceleration, latent-domain multi-resolution methods, super-resolution, and repainting; Appendix B records baseline settings, guidance, and timing scope. Appendix C measures greater text-to-image attention mass at 512 than 1024 resolution in three sampled Qwen-Image layers and reports 5.08s for LR10 HR1 versus 10.96s for HR11. Its super-resolution and frequency-band analyses find short refinement less able to recover from low-frequency perturbations than from the tested high-frequency perturbations.
- Appendix C derives the noise threshold under a local anisotropic-Gaussian prior and reports a high-frequency threshold of approximately 0.075 from 18 clean latents. It separately gives a distribution-level Gaussian-smoothing bound, which concerns overall distribution alignment rather than the weaker direction-wise resampling condition. Refinement-step and seed-sensitivity tests support the measured structure–detail division, subject to the local-prior and high-frequency-residual assumptions.
- Appendices D and E extend evaluation to FLUX.2 Klein, Z-Image, and reduced-step variants, and report Qwen-Image runtime of 49.32s natively versus 4.77s for MrFlow. Transfer is not uniform in the measured metrics: Z-Image at (12, 1)×2 reaches 10.8×, while OneIG-En/OneIG-Zh fall from 0.56/0.52 to 0.51/0.46. Appendix F provides method comparisons, additional image examples, and their prompts, and notes that PDF compression can obscure fine details.
Brief Thoughts
End-to-end timing accounts for the extra VAE and super-resolution work, while the similar automatic scores across super-resolution choices make the local visual comparisons relevant. The noise analysis applies when super-resolution preserves structure and leaves locally plausible, predominantly high-frequency errors; the Z-Image result cautions against assuming equal quality retention across backbones. Table 1 and Table 2 also report different DPG values, 81.65 and 81.08, for the same FLUX.1-dev MrFlow (12, 1) setting.