Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing | Summary
21 Jul 2026 | Paper Review Text-to-Image Generation Instruction-Based Image Editing Diffusion Transformers Flow MatchingContents
- Summary
- 3 Mage-Flow Stack
- 4 Data Collection and Curation
- 5 Training
- 6 Experiments
- Appendix
- Brief Thoughts
This article explains the key points of Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing.
- 2026-07-21 (arXiv)
- Zhang, Xinjie, Zhang, Peng, Zheng, Shicheng, Guo, Jinghao, Jia, Zhaoyang, Shen, Yifei, Guo, Xun, Luo, Yuxuan, Li, Jiahao, Xie, Wenxuan, et al.
- Microsoft
- Paper
- Project Page
Summary
- Mage-Flow is a 4B-scale foundation-model stack for native-resolution text-to-image generation and instruction-based image editing. It combines Mage-VAE, a lightweight latent tokenizer, with a Native-Resolution Multimodal Diffusion Transformer trained by rectified flow matching; reported 1024² single-A100 Turbo latencies are 0.59 s for generation and 1.02 s for editing.
Source-centric examples cover background, color, count, font, text, structural-map, and pose-related transformations.
This gallery extends the demonstrated edits to material, style, removal, retouching, weather, depth, viewpoint, grayscale, extraction, and Canny-edge transformations.
The scatterplots compare benchmark score, end-to-end A100 latency, and peak memory for representative generation and editing systems. FLUX.2-dev is evaluated on two A100 GPUs, unlike the other entries.
3 Mage-Flow Stack
The stack couples Mage-VAE with a 4B NR-MMDiT and training infrastructure intended to reduce tokenization, batching, and memory-bound systems overhead. It supports Base, Diffusion-NFT-aligned, and four-step Turbo variants for generation and editing.
The diagram shows variable-length text and image sequences packed for NR-MMDiT, clarifying that native resolution is retained through packed sequences rather than fixed buckets.
The figure connects FLUX.2-VAE anchor latents with Mage-VAE's symmetric one-step encoder-decoder and its three-stage training procedure.
3.1.1 Architecture
Mage-VAE uses a fully convolutional one-step pixel-diffusion decoder composed of convolutional diffusion blocks and a decoupled pixel-diffusion head. Its encoder is an architectural dual: patch embedding followed by convolutional diffusion blocks generates latents conditioned on pixels, avoiding global attention in both directions.
3.1.2 KL regularization with an anchor latent distribution
Rather than regularizing against a standard Gaussian prior, Mage-VAE aligns its posterior with FLUX.2-VAE anchor latents. It internalizes FLUX.2-VAE’s 2× latent patchification and directly produces 16×-downsampled, 128-channel Transformer-ready latents.
3.1.3 Training
Mage-VAE training has three stages: separate multi-step flow-matching pre-training of encoder and decoder, one-step decoder distillation using reconstruction, DMD, and DINOv2-projected GAN losses, and joint encoder-decoder fine-tuning with anchor-latent KL regularization. The final tokenizer is designed to retain perceptual reconstruction while remaining compatible with downstream generation.
3.1.4 Evaluation
On CLIC 2020, Mage-VAE reports 36.61 PSNR, 0.9450 SSIM, and 0.0148 LPIPS; on FFHQ validation at 1024 × 1024, it reports 40.67 PSNR, 0.9708 SSIM, and 0.0107 LPIPS. Compared with FLUX.2-VAE, its encoder and decoder complexity is 173 and 215 kMACs/pixel versus 2134 and 4798 kMACs/pixel, while cross-tokenizer tests show broadly similar downstream scores.
Mage-VAE has 173 versus 2134 encoder kMACs/pixel and 215 versus 4798 decoder kMACs/pixel relative to FLUX.2-VAE, while reconstruction metrics remain close.
The curves show Mage-VAE retaining substantially lower latency at increasing resolution; several compared tokenizers become extremely slow or run out of memory at 4096².
Swapping Mage-VAE and FLUX.2-VAE while holding a downstream backbone fixed produces similar results across the reported generation and editing metrics.
3.2.1 Model Architecture and Native Packing
NR-MMDiT concatenates text and image tokens for joint self-attention while retaining modality-specific normalization and projections. Native-resolution packing flattens arbitrary-resolution latent grids and variable-length text under a fixed token budget using FlashAttention variable-length kernels and per-sample 2D RoPE; packed classifier-free guidance reports 1.09×–1.15× inference speedups.
Packing conditional and unconditional branches reduces reported inference time while preserving the specified denoising trajectory, with gains from 1.09× to 1.15×.
3.2.2 Conditioning and rectified-flow training
A frozen Qwen3-VL-4B-Instruct provides text embeddings, and Mage-VAE maps images to 16×-reduced, 128-channel latents for rectified-flow training. For editing, the same backbone processes instruction embeddings, source-image latents, and noisy target latents; frame-aware 3D RoPE separates source and target tokens, and loss is computed only for targets.
3.3 Training Infrastructure
The implementation fuses memory-bound operation chains in Mage-VAE, Qwen3-VL, and NR-MMDiT, including normalization, modulation, RoPE, gating, activation, and residual operations. On one 8-GPU NVIDIA B200 node, the full configuration increases MFU from 13.88% to 29.28%, lowers peak per-GPU memory from 175.45 GB to 141.44 GB, and reduces step time from 1.9285 s to 0.7775 s, a reported 2.48× speedup.
The ablation separates the tokenizer and fusion effects: Mage-VAE alone gives 1.41× speedup, and the complete fused configuration reaches 2.48× with lower memory.
4 Data Collection and Curation
The data strategy uses complementary pipelines for text-to-image pairs and source-instruction-target editing triples. Both pipelines target visual quality, safety, semantic alignment, diversity, and coverage of capabilities underrepresented in raw web data.
4.1 Text-to-Image Data Collection and Filtering
Starting from roughly 10B open-source image-text pairs, the generation pipeline applies file and content filtering, SSCD/FAISS-based deduplication, Qwen3-VL-32B-Instruct multi-granularity recaptioning, and concept-aware synthesis, retaining approximately 1.3B pairs. Later stages tighten resolution, watermark, aesthetic, OCR, and safety thresholds while supplementing long-text rendering, rare concepts, layouts, and styles.
The pipeline traces the 1.3B-pair corpus from filtering and deduplication through multi-level recaptioning and concept-aware synthesis.
The listed thresholds become stricter at later stages, especially for resolution, watermark probability, and aesthetic score.
The sunbursts show generation data concentrated in Object & Products and Scene & Place, while editing data is spread across semantic, attribute, text, object, style, and low-level operations.
4.2 Image-Editing Data Collection and Filtering
The editing pool begins with roughly 90M triples: about 50M from open-source datasets and 40M synthesized in-house. Three Qwen3.5-9B experts judge instruction execution, preservation, and plausibility; majority voting retains roughly 45M triples, which are mapped to 19 edit categories and rebalanced.
The diagram shows three VLM experts filtering editing triples by majority vote before taxonomy assignment and category balancing.
5 Training
Generation and editing share the Mage-VAE latent space, the 4B NR-MMDiT, and rectified-flow training. The training tree first produces Mage-Flow-Base, then branches into aligned and Turbo generation models and source-conditioned Base, aligned, and Turbo editing models.
The training tree distinguishes progressive generation training, editing continuation training, Diffusion-NFT post-training, and four-step distillation.
5.1 Pre-training and Supervised Fine-tuning
Generation training progresses through 1.2B pairs at fixed 256 × 256, 600M higher-quality samples in a 512-pixel native-aspect-ratio regime, 300M cleaner samples in a 1024-pixel regime, and SFT on 150M curated pairs. Editing initializes from Mage-Flow-Base, then trains on 35M editing triples plus 35M generation pairs and on 20M editing triples plus 10M generation pairs to preserve the open-ended generation prior during adaptation.
The visual presents a unified image-and-text-conditioned editor producing semantic edits, restorations, stylization, camera changes, and structural outputs.
5.2 Diffusion-NFT Post-training
Diffusion-NFT performs negative-aware online fine-tuning on the forward process without likelihood estimation. Generation uses approximately 20K prompts routed to OCR, aesthetic, or semantic evaluators; editing uses a 4:1 editing-to-generation update ratio, approximately 30K editing prompts, and RationalRewards for instruction adherence, preservation, plausibility, and text rendering.
5.3 Few-step Distillation
Turbo variants distill aligned teachers into four-step students using Decoupled DMD and adversarial perceptual guidance based on frozen DINOv2 and CLIP features. The reported ablation shows benefits for most generation and text-editing metrics, while general-editing changes are mixed; adding generation data increases Turbo ImgEdit from 4.20 to 4.38 without uniformly improving GEdit.
The schematic separates Decoupled DMD's classifier-free augmentation and distribution-matching signals from feature-space adversarial guidance.
Adversarial guidance improves many reported generation and text-editing scores, but the reported general-editing changes are benchmark-dependent.
Generation-data mixing raises Turbo ImgEdit from 4.20 to 4.38, whereas English and Chinese GEdit changes are not uniform.
The two sunbursts document category coverage in generation and editing distillation data, including people, scenes, products, text, controls, and complex edits.
6 Experiments
Experiments compare Base, aligned, and Turbo variants for generation and editing through quantitative benchmarks and qualitative galleries. The evaluation addresses quality, prompt following, text rendering, editing accuracy, latency, and memory rather than a single benchmark.
6.1 Experimental Setup
Generation is evaluated at native 1024² using GenEval, DPG-Bench, TIIF-Bench, OneIG, CVTG-2K, and LongText; Base, aligned, and Turbo variants use 30, 20, and 4 steps. Editing is evaluated with ImgEdit-Bench, GEdit-Bench, and TextEdit-Bench; Base and aligned models use 30 steps, while Turbo uses 4.
6.2 Quantitative Results
Mage-Flow reports a 0.90 GenEval score at 20 steps and 0.88 at four steps; its CVTG-2K averages are 0.887 and 0.873, respectively. Mage-Flow-Edit-Turbo reports 4.38 on ImgEdit, 8.271/8.264 on GEdit-Bench English/Chinese, and 12.77/15.41 on TextEdit-Bench Synthetic/Real-World; the paper identifies precise layout preservation and complex text replacement as remaining limitations.
Mage-Flow reports 0.90 GenEval at 20 steps and Turbo reports 0.88 at four steps, while relative performance varies across the eight benchmarks.
The breakdown reports strong Mage-Flow scores on single-object, two-object, counting, and positional GenEval categories, while DPG-Bench is competitive rather than uniformly leading.
TIIF-Bench separates short and long prompts across basic, advanced, and designer tasks, showing different behavior by task category and step budget.
Mage-Flow reports a 0.887 CVTG-2K average and Turbo 0.873; scores are broken down by two to five text regions.
OneIG reports strong alignment and text subscores for Mage-Flow variants but lower diversity and lower Chinese overall scores than some competitors.
The summary table compares editing systems at their reported step counts and shows the quality-efficiency trade-off among the 4B Mage variants and larger models.
Category-level ImgEdit results show Turbo at 4.38 overall, with substantial variation across extraction, hybrid, style, and other edit classes.
Mage-Flow-Edit-Turbo reports GEdit overall scores of 8.271 in English and 8.264 in Chinese, alongside semantic-consistency and perceptual-quality components.
Among Mage variants, the aligned 30-step editor is strongest on TextEdit-Bench, reporting 14.14 Synthetic and 16.26 Real-World; Turbo scores 12.77 and 15.41.
The gallery displays uncropped native-aspect-ratio cuisine and landmark generations across close-up and wide scenes.
Examples cover wildlife, scenery, sculptures, and surreal composites to illustrate subject and style breadth.
The portrait gallery shows people across settings, ages, cultural dress, and lighting conditions.
The examples focus on multi-line English text in packaging, posters, covers, and signs, complementing the text-rendering evaluations.
Chinese and bilingual signage, packaging, and cover layouts provide qualitative examples of multilingual text rendering.
Paired source-output examples show localized text replacement, object counting, addition, removal, replacement, and extraction.
The panel demonstrates background replacement, zoom, viewpoint, scale, action, and composite scene changes.
Examples show tone, style, material, and color transformations for landscapes, food, architecture, and objects.
The gallery includes virtual try-on, portrait retouching, hairstyle edits, meme composition, and multi-instruction transformations.
Paired examples cover bidirectional lens flare, rain, haze, blur, grayscale, colorization, exposure, and enhancement tasks.
The examples cover sketches, pose, Canny and HED maps, depth, segmentation, normal maps, and image reconstruction from controls.
Appendix
- Appendix A details Mage-VAE objectives and joint optimization; Appendix B reports 149.6/375.0 ms encoding/decoding at 4096² on an 80GB A100. The appendices also specify reward evaluators and describe scientific-diagram fine-tuning, where Mage-Flow-SciForma reports 61.61 on SciFormaBench-2K versus 40.80 for zero-shot Mage-Flow-Base.
Brief Thoughts
The paper’s efficiency case depends on co-designing the tokenizer, native-resolution packing, and fused kernels rather than scaling the backbone alone. Its comparisons and several evaluation benchmarks rely on reported or model-judged scores, and the authors note incomplete multilingual long-text rendering, limited multi-image editing coverage, and non-uniform distillation gains for general editing.