Gorio Tech Blog search

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing | Summary

|

Contents

This article explains the key points of Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing.

  • 2026-07-21 (arXiv)
  • Zhang, Xinjie, Zhang, Peng, Zheng, Shicheng, Guo, Jinghao, Jia, Zhaoyang, Shen, Yifei, Guo, Xun, Luo, Yuxuan, Li, Jiahao, Xie, Wenxuan, et al.
  • Microsoft
  • Paper
  • Project Page

Read this article in Korean


Summary

  • Mage-Flow is a 4B-scale foundation-model stack for native-resolution text-to-image generation and instruction-based image editing. It combines Mage-VAE, a lightweight latent tokenizer, with a Native-Resolution Multimodal Diffusion Transformer trained by rectified flow matching; reported 1024² single-A100 Turbo latencies are 0.59 s for generation and 1.02 s for editing.

Source-centric examples cover background, color, count, font, text, structural-map, and pose-related transformations.

Source-centric examples of bidirectional background, color, count, font, effect, object, text, segmentation, normal-map, action, HED-edge, and pose edits.
Source-centric examples of bidirectional background, color, count, font, effect, object, text, segmentation, normal-map, action, HED-edge, and pose edits.

This gallery extends the demonstrated edits to material, style, removal, retouching, weather, depth, viewpoint, grayscale, extraction, and Canny-edge transformations.

Additional source-centric examples of material, style, removal, retouching, weather, depth, tone, viewpoint, grayscale, extraction, Canny-edge, and composite edits.
Additional source-centric examples of material, style, removal, retouching, weather, depth, tone, viewpoint, grayscale, extraction, Canny-edge, and composite edits.

The scatterplots compare benchmark score, end-to-end A100 latency, and peak memory for representative generation and editing systems. FLUX.2-dev is evaluated on two A100 GPUs, unlike the other entries.

Quality, latency, and peak-memory comparison for representative text-to-image and image-editing models on NVIDIA A100 GPUs.
Quality, latency, and peak-memory comparison for representative text-to-image and image-editing models on NVIDIA A100 GPUs.

3 Mage-Flow Stack

The stack couples Mage-VAE with a 4B NR-MMDiT and training infrastructure intended to reduce tokenization, batching, and memory-bound systems overhead. It supports Base, Diffusion-NFT-aligned, and four-step Turbo variants for generation and editing.

The diagram shows variable-length text and image sequences packed for NR-MMDiT, clarifying that native resolution is retained through packed sequences rather than fixed buckets.

Mage-Flow architecture with Qwen3-VL conditioning, Mage-VAE tokenization, packed native-resolution sequences, and NR-MMDiT blocks.
Mage-Flow architecture with Qwen3-VL conditioning, Mage-VAE tokenization, packed native-resolution sequences, and NR-MMDiT blocks.

The figure connects FLUX.2-VAE anchor latents with Mage-VAE's symmetric one-step encoder-decoder and its three-stage training procedure.

Mage-VAE anchor-latent design, lightweight architecture, and diffusion pre-training, decoder-distillation, and joint-distillation stages.
Mage-VAE anchor-latent design, lightweight architecture, and diffusion pre-training, decoder-distillation, and joint-distillation stages.

3.1.1 Architecture

Mage-VAE uses a fully convolutional one-step pixel-diffusion decoder composed of convolutional diffusion blocks and a decoupled pixel-diffusion head. Its encoder is an architectural dual: patch embedding followed by convolutional diffusion blocks generates latents conditioned on pixels, avoiding global attention in both directions.

3.1.2 KL regularization with an anchor latent distribution

Rather than regularizing against a standard Gaussian prior, Mage-VAE aligns its posterior with FLUX.2-VAE anchor latents. It internalizes FLUX.2-VAE’s 2× latent patchification and directly produces 16×-downsampled, 128-channel Transformer-ready latents.

3.1.3 Training

Mage-VAE training has three stages: separate multi-step flow-matching pre-training of encoder and decoder, one-step decoder distillation using reconstruction, DMD, and DINOv2-projected GAN losses, and joint encoder-decoder fine-tuning with anchor-latent KL regularization. The final tokenizer is designed to retain perceptual reconstruction while remaining compatible with downstream generation.

3.1.4 Evaluation

On CLIC 2020, Mage-VAE reports 36.61 PSNR, 0.9450 SSIM, and 0.0148 LPIPS; on FFHQ validation at 1024 × 1024, it reports 40.67 PSNR, 0.9708 SSIM, and 0.0107 LPIPS. Compared with FLUX.2-VAE, its encoder and decoder complexity is 173 and 215 kMACs/pixel versus 2134 and 4798 kMACs/pixel, while cross-tokenizer tests show broadly similar downstream scores.

Mage-VAE has 173 versus 2134 encoder kMACs/pixel and 215 versus 4798 decoder kMACs/pixel relative to FLUX.2-VAE, while reconstruction metrics remain close.

Mage-VAE parameter counts, encoding and decoding complexity, and reconstruction metrics on CLIC 2020 and FFHQ.
Mage-VAE parameter counts, encoding and decoding complexity, and reconstruction metrics on CLIC 2020 and FFHQ.

The curves show Mage-VAE retaining substantially lower latency at increasing resolution; several compared tokenizers become extremely slow or run out of memory at 4096².

Encoding and decoding latency and memory behavior across resolutions on an 80GB A100 GPU.
Encoding and decoding latency and memory behavior across resolutions on an 80GB A100 GPU.

Swapping Mage-VAE and FLUX.2-VAE while holding a downstream backbone fixed produces similar results across the reported generation and editing metrics.

Cross-tokenizer compatibility results for Mage-VAE and FLUX.2-VAE in generation and instruction-based editing backbones.
Cross-tokenizer compatibility results for Mage-VAE and FLUX.2-VAE in generation and instruction-based editing backbones.

3.2.1 Model Architecture and Native Packing

NR-MMDiT concatenates text and image tokens for joint self-attention while retaining modality-specific normalization and projections. Native-resolution packing flattens arbitrary-resolution latent grids and variable-length text under a fixed token budget using FlashAttention variable-length kernels and per-sample 2D RoPE; packed classifier-free guidance reports 1.09×–1.15× inference speedups.

Packing conditional and unconditional branches reduces reported inference time while preserving the specified denoising trajectory, with gains from 1.09× to 1.15×.

Inference times and speedups from packed classifier-free-guidance evaluation on one NVIDIA A100 GPU.
Inference times and speedups from packed classifier-free-guidance evaluation on one NVIDIA A100 GPU.

3.2.2 Conditioning and rectified-flow training

A frozen Qwen3-VL-4B-Instruct provides text embeddings, and Mage-VAE maps images to 16×-reduced, 128-channel latents for rectified-flow training. For editing, the same backbone processes instruction embeddings, source-image latents, and noisy target latents; frame-aware 3D RoPE separates source and target tokens, and loss is computed only for targets.

3.3 Training Infrastructure

The implementation fuses memory-bound operation chains in Mage-VAE, Qwen3-VL, and NR-MMDiT, including normalization, modulation, RoPE, gating, activation, and residual operations. On one 8-GPU NVIDIA B200 node, the full configuration increases MFU from 13.88% to 29.28%, lowers peak per-GPU memory from 175.45 GB to 141.44 GB, and reduces step time from 1.9285 s to 0.7775 s, a reported 2.48× speedup.

The ablation separates the tokenizer and fusion effects: Mage-VAE alone gives 1.41× speedup, and the complete fused configuration reaches 2.48× with lower memory.

Training-efficiency ablation for Mage-VAE substitution and fused CUDA kernels on an 8-GPU NVIDIA B200 node.
Training-efficiency ablation for Mage-VAE substitution and fused CUDA kernels on an 8-GPU NVIDIA B200 node.

4 Data Collection and Curation

The data strategy uses complementary pipelines for text-to-image pairs and source-instruction-target editing triples. Both pipelines target visual quality, safety, semantic alignment, diversity, and coverage of capabilities underrepresented in raw web data.

4.1 Text-to-Image Data Collection and Filtering

Starting from roughly 10B open-source image-text pairs, the generation pipeline applies file and content filtering, SSCD/FAISS-based deduplication, Qwen3-VL-32B-Instruct multi-granularity recaptioning, and concept-aware synthesis, retaining approximately 1.3B pairs. Later stages tighten resolution, watermark, aesthetic, OCR, and safety thresholds while supplementing long-text rendering, rare concepts, layouts, and styles.

The pipeline traces the 1.3B-pair corpus from filtering and deduplication through multi-level recaptioning and concept-aware synthesis.

Text-to-image filtering, deduplication, multi-granularity captioning, and concept-aware synthesis pipeline.
Text-to-image filtering, deduplication, multi-granularity captioning, and concept-aware synthesis pipeline.

The listed thresholds become stricter at later stages, especially for resolution, watermark probability, and aesthetic score.

Sample-level filtering thresholds for 256², 512², 1024², and supervised fine-tuning data stages.
Sample-level filtering thresholds for 256², 512², 1024², and supervised fine-tuning data stages.

The sunbursts show generation data concentrated in Object & Products and Scene & Place, while editing data is spread across semantic, attribute, text, object, style, and low-level operations.

Concept and edit-category composition of generation and editing pre-training mixtures.
Concept and edit-category composition of generation and editing pre-training mixtures.

4.2 Image-Editing Data Collection and Filtering

The editing pool begins with roughly 90M triples: about 50M from open-source datasets and 40M synthesized in-house. Three Qwen3.5-9B experts judge instruction execution, preservation, and plausibility; majority voting retains roughly 45M triples, which are mapped to 19 edit categories and rebalanced.

The diagram shows three VLM experts filtering editing triples by majority vote before taxonomy assignment and category balancing.

Editing-triple collection, VLM-majority filtering, edit-type tagging, and category-balancing pipeline.
Editing-triple collection, VLM-majority filtering, edit-type tagging, and category-balancing pipeline.

5 Training

Generation and editing share the Mage-VAE latent space, the 4B NR-MMDiT, and rectified-flow training. The training tree first produces Mage-Flow-Base, then branches into aligned and Turbo generation models and source-conditioned Base, aligned, and Turbo editing models.

The training tree distinguishes progressive generation training, editing continuation training, Diffusion-NFT post-training, and four-step distillation.

Training pipeline from progressive generation pre-training to Base, aligned, and Turbo checkpoints for generation and editing.
Training pipeline from progressive generation pre-training to Base, aligned, and Turbo checkpoints for generation and editing.

5.1 Pre-training and Supervised Fine-tuning

Generation training progresses through 1.2B pairs at fixed 256 × 256, 600M higher-quality samples in a 512-pixel native-aspect-ratio regime, 300M cleaner samples in a 1024-pixel regime, and SFT on 150M curated pairs. Editing initializes from Mage-Flow-Base, then trains on 35M editing triples plus 35M generation pairs and on 20M editing triples plus 10M generation pairs to preserve the open-ended generation prior during adaptation.

The visual presents a unified image-and-text-conditioned editor producing semantic edits, restorations, stylization, camera changes, and structural outputs.

Unified Mage-Flow-Edit capabilities and representative outputs for semantic, appearance, restoration, and structure-aware tasks.
Unified Mage-Flow-Edit capabilities and representative outputs for semantic, appearance, restoration, and structure-aware tasks.

5.2 Diffusion-NFT Post-training

Diffusion-NFT performs negative-aware online fine-tuning on the forward process without likelihood estimation. Generation uses approximately 20K prompts routed to OCR, aesthetic, or semantic evaluators; editing uses a 4:1 editing-to-generation update ratio, approximately 30K editing prompts, and RationalRewards for instruction adherence, preservation, plausibility, and text rendering.

5.3 Few-step Distillation

Turbo variants distill aligned teachers into four-step students using Decoupled DMD and adversarial perceptual guidance based on frozen DINOv2 and CLIP features. The reported ablation shows benefits for most generation and text-editing metrics, while general-editing changes are mixed; adding generation data increases Turbo ImgEdit from 4.20 to 4.38 without uniformly improving GEdit.

The schematic separates Decoupled DMD's classifier-free augmentation and distribution-matching signals from feature-space adversarial guidance.

Few-step distillation framework combining Decoupled DMD with DINOv2- and CLIP-based adversarial perceptual guidance.
Few-step distillation framework combining Decoupled DMD with DINOv2- and CLIP-based adversarial perceptual guidance.

Adversarial guidance improves many reported generation and text-editing scores, but the reported general-editing changes are benchmark-dependent.

Effect of adversarial perceptual guidance on four-step generation and editing distillation.
Effect of adversarial perceptual guidance on four-step generation and editing distillation.

Generation-data mixing raises Turbo ImgEdit from 4.20 to 4.38, whereas English and Chinese GEdit changes are not uniform.

Effect of mixing generation data during full-step and four-step image-editing training.
Effect of mixing generation data during full-step and four-step image-editing training.

The two sunbursts document category coverage in generation and editing distillation data, including people, scenes, products, text, controls, and complex edits.

Category and sub-category composition of generation and editing distillation datasets.
Category and sub-category composition of generation and editing distillation datasets.

6 Experiments

Experiments compare Base, aligned, and Turbo variants for generation and editing through quantitative benchmarks and qualitative galleries. The evaluation addresses quality, prompt following, text rendering, editing accuracy, latency, and memory rather than a single benchmark.

6.1 Experimental Setup

Generation is evaluated at native 1024² using GenEval, DPG-Bench, TIIF-Bench, OneIG, CVTG-2K, and LongText; Base, aligned, and Turbo variants use 30, 20, and 4 steps. Editing is evaluated with ImgEdit-Bench, GEdit-Bench, and TextEdit-Bench; Base and aligned models use 30 steps, while Turbo uses 4.

6.2 Quantitative Results

Mage-Flow reports a 0.90 GenEval score at 20 steps and 0.88 at four steps; its CVTG-2K averages are 0.887 and 0.873, respectively. Mage-Flow-Edit-Turbo reports 4.38 on ImgEdit, 8.271/8.264 on GEdit-Bench English/Chinese, and 12.77/15.41 on TextEdit-Bench Synthetic/Real-World; the paper identifies precise layout preservation and complex text replacement as remaining limitations.

Mage-Flow reports 0.90 GenEval at 20 steps and Turbo reports 0.88 at four steps, while relative performance varies across the eight benchmarks.

Text-to-image generation results across GenEval, DPG, TIIF, CVTG-2K, OneIG, and LongText benchmarks.
Text-to-image generation results across GenEval, DPG, TIIF, CVTG-2K, OneIG, and LongText benchmarks.

The breakdown reports strong Mage-Flow scores on single-object, two-object, counting, and positional GenEval categories, while DPG-Bench is competitive rather than uniformly leading.

Detailed GenEval and DPG-Bench results for closed-source, unified, specialist, and Mage-Flow models.
Detailed GenEval and DPG-Bench results for closed-source, unified, specialist, and Mage-Flow models.

TIIF-Bench separates short and long prompts across basic, advanced, and designer tasks, showing different behavior by task category and step budget.

TIIF-Bench testmini results across instruction-following and designer-oriented prompt categories.
TIIF-Bench testmini results across instruction-following and designer-oriented prompt categories.

Mage-Flow reports a 0.887 CVTG-2K average and Turbo 0.873; scores are broken down by two to five text regions.

Multi-region English text-rendering results on CVTG-2K.
Multi-region English text-rendering results on CVTG-2K.

OneIG reports strong alignment and text subscores for Mage-Flow variants but lower diversity and lower Chinese overall scores than some competitors.

Fine-grained generation results on OneIG English and Chinese splits.
Fine-grained generation results on OneIG English and Chinese splits.

The summary table compares editing systems at their reported step counts and shows the quality-efficiency trade-off among the 4B Mage variants and larger models.

Instruction-based image-editing results on ImgEdit-Bench, GEdit-Bench, and TextEdit-Bench.
Instruction-based image-editing results on ImgEdit-Bench, GEdit-Bench, and TextEdit-Bench.

Category-level ImgEdit results show Turbo at 4.38 overall, with substantial variation across extraction, hybrid, style, and other edit classes.

Category-wise ImgEdit-Bench results for nine editing skills.
Category-wise ImgEdit-Bench results for nine editing skills.

Mage-Flow-Edit-Turbo reports GEdit overall scores of 8.271 in English and 8.264 in Chinese, alongside semantic-consistency and perceptual-quality components.

Semantic consistency, perceptual quality, and overall GEdit-Bench scores for English and Chinese.
Semantic consistency, perceptual quality, and overall GEdit-Bench scores for English and Chinese.

Among Mage variants, the aligned 30-step editor is strongest on TextEdit-Bench, reporting 14.14 Synthetic and 16.26 Real-World; Turbo scores 12.77 and 15.41.

TextEdit-Bench instruction following, text accuracy, visual consistency, layout preservation, and semantic expectation scores.
TextEdit-Bench instruction following, text accuracy, visual consistency, layout preservation, and semantic expectation scores.

The gallery displays uncropped native-aspect-ratio cuisine and landmark generations across close-up and wide scenes.

Mage-Flow-Base native-resolution gallery of cuisine and world landmarks.
Mage-Flow-Base native-resolution gallery of cuisine and world landmarks.

Examples cover wildlife, scenery, sculptures, and surreal composites to illustrate subject and style breadth.

Mage-Flow gallery of general and imaginative concepts.
Mage-Flow gallery of general and imaginative concepts.

The portrait gallery shows people across settings, ages, cultural dress, and lighting conditions.

Mage-Flow gallery of portraits and people.
Mage-Flow gallery of portraits and people.

The examples focus on multi-line English text in packaging, posters, covers, and signs, complementing the text-rendering evaluations.

Mage-Flow gallery of English text rendering on packaging, covers, posters, and signage.
Mage-Flow gallery of English text rendering on packaging, covers, posters, and signage.

Chinese and bilingual signage, packaging, and cover layouts provide qualitative examples of multilingual text rendering.

Mage-Flow gallery of Chinese and multilingual text rendering.
Mage-Flow gallery of Chinese and multilingual text rendering.

Paired source-output examples show localized text replacement, object counting, addition, removal, replacement, and extraction.

Mage-Flow-Edit qualitative gallery of localized content and object editing.
Mage-Flow-Edit qualitative gallery of localized content and object editing.

The panel demonstrates background replacement, zoom, viewpoint, scale, action, and composite scene changes.

Mage-Flow-Edit qualitative gallery of scene, subject, and camera transformations.
Mage-Flow-Edit qualitative gallery of scene, subject, and camera transformations.

Examples show tone, style, material, and color transformations for landscapes, food, architecture, and objects.

Mage-Flow-Edit qualitative gallery of appearance and artistic rendering.
Mage-Flow-Edit qualitative gallery of appearance and artistic rendering.

The gallery includes virtual try-on, portrait retouching, hairstyle edits, meme composition, and multi-instruction transformations.

Mage-Flow-Edit qualitative gallery of human-centered and creative editing.
Mage-Flow-Edit qualitative gallery of human-centered and creative editing.

Paired examples cover bidirectional lens flare, rain, haze, blur, grayscale, colorization, exposure, and enhancement tasks.

Mage-Flow-Edit qualitative gallery of bidirectional image degradation and restoration.
Mage-Flow-Edit qualitative gallery of bidirectional image degradation and restoration.

The examples cover sketches, pose, Canny and HED maps, depth, segmentation, normal maps, and image reconstruction from controls.

Mage-Flow-Edit qualitative gallery of low-level vision and conditional reconstruction.
Mage-Flow-Edit qualitative gallery of low-level vision and conditional reconstruction.

Appendix

  • Appendix A details Mage-VAE objectives and joint optimization; Appendix B reports 149.6/375.0 ms encoding/decoding at 4096² on an 80GB A100. The appendices also specify reward evaluators and describe scientific-diagram fine-tuning, where Mage-Flow-SciForma reports 61.61 on SciFormaBench-2K versus 40.80 for zero-shot Mage-Flow-Base.

Brief Thoughts

The paper’s efficiency case depends on co-designing the tokenizer, native-resolution packing, and fused kernels rather than scaling the backbone alone. Its comparisons and several evaluation benchmarks rely on reported or model-judged scores, and the authors note incomplete multilingual long-text rendering, limited multi-image editing coverage, and non-uniform distillation gains for general editing.