Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing 요약 설명
21 Jul 2026 | Paper Review Text-to-Image Generation Instruction-Based Image Editing Diffusion Transformers Flow Matching목차
이번 글에서는 Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing 논문의 핵심 포인트만 간단히 정리한다.
- 2026년 7월 21일(Arxiv)
- Zhang, Xinjie, Zhang, Peng, Zheng, Shicheng, Guo, Jinghao, Jia, Zhaoyang, Shen, Yifei, Guo, Xun, Luo, Yuxuan, Li, Jiahao, Xie, Wenxuan, et al.
- Microsoft
- 논문 링크
- Project Page
요약
- Mage-Flow는 native-resolution 텍스트-이미지 생성과 instruction-based 이미지 편집을 위한 4B-scale foundation-model 스택이다. 경량 latent tokenizer인 Mage-VAE와 rectified flow matching으로 학습한 Native-Resolution Multimodal Diffusion Transformer를 결합하며, 보고된 1024² 단일 A100 Turbo 지연 시간은 생성 0.59 s, 편집 1.02 s다.
Source-centric example은 background, color, count, font, text, structural map, pose 관련 변환을 다룬다.
이 gallery는 material, style, removal, retouching, weather, depth, viewpoint, grayscale, extraction, Canny-edge 변환까지 제시한 편집을 확장한다.
Scatterplot은 대표적인 생성 및 편집 시스템의 benchmark score, end-to-end A100 latency, peak memory를 비교한다. FLUX.2-dev는 다른 항목과 달리 A100 GPU 2개에서 평가한다.
3 Mage-Flow Stack
이 스택은 Mage-VAE, 4B NR-MMDiT, 그리고 tokenization, batching, memory-bound 시스템 오버헤드를 줄이기 위한 학습 인프라를 결합한다. 생성과 편집에 Base, Diffusion-NFT-aligned, four-step Turbo 변형을 지원한다.
도식은 NR-MMDiT를 위해 variable-length text와 image sequence를 packing하는 방식을 보여주며, native resolution이 fixed bucket이 아니라 packed sequence로 유지됨을 명확히 한다.
그림은 FLUX.2-VAE anchor latent를 Mage-VAE의 대칭적 one-step encoder-decoder 및 3단계 학습 절차와 연결한다.
3.1.1 Architecture
Mage-VAE는 convolutional diffusion block과 분리된 pixel-diffusion head로 구성한 fully convolutional one-step pixel-diffusion decoder를 사용한다. encoder는 이에 대응하는 구조로, patch embedding 뒤에 convolutional diffusion block을 적용해 pixel 조건부 latent를 생성하며, 양방향 모두에서 global attention을 피한다.
3.1.2 KL regularization with an anchor latent distribution
Mage-VAE는 표준 Gaussian prior에 정규화하는 대신 posterior를 FLUX.2-VAE anchor latent에 정렬한다. FLUX.2-VAE의 2× latent patchification을 내부화하고, 16× downsampled 128-channel Transformer-ready latent를 직접 생성한다.
3.1.3 Training
Mage-VAE 학습은 3단계로 구성한다. 먼저 encoder와 decoder를 별도로 multi-step flow-matching pre-training하고, reconstruction, DMD, DINOv2-projected GAN loss로 one-step decoder distillation을 수행한 뒤, anchor-latent KL regularization으로 encoder-decoder를 jointly fine-tuning한다. 최종 tokenizer는 지각적 재구성 성능을 유지하면서 downstream 생성과 호환되도록 설계했다. GAN에 대해서는 이 글을 참조하라.
3.1.4 Evaluation
CLIC 2020에서 Mage-VAE는 36.61 PSNR, 0.9450 SSIM, 0.0148 LPIPS를 보고하며, 1024 × 1024 FFHQ validation에서는 40.67 PSNR, 0.9708 SSIM, 0.0107 LPIPS를 보고한다. FLUX.2-VAE와 비교하면 encoder와 decoder 복잡도는 각각 2134와 4798 kMACs/pixel 대비 173과 215 kMACs/pixel이며, cross-tokenizer 실험에서는 전반적으로 유사한 downstream score를 보인다.
Mage-VAE는 FLUX.2-VAE 대비 173 대 2134 encoder kMACs/pixel, 215 대 4798 decoder kMACs/pixel을 사용하면서 재구성 metric은 근접하게 유지한다.
곡선은 해상도가 증가할 때 Mage-VAE가 훨씬 낮은 latency를 유지함을 보여주며, 비교 tokenizer 일부는 4096²에서 매우 느려지거나 memory 부족이 발생한다.
Downstream backbone을 고정한 채 Mage-VAE와 FLUX.2-VAE를 교체하면 보고된 생성 및 편집 metric에서 유사한 결과가 나온다.
3.2.1 Model Architecture and Native Packing
NR-MMDiT는 text token과 image token을 연결해 joint self-attention을 수행하면서 modality-specific normalization과 projection을 유지한다. Native-resolution packing은 FlashAttention variable-length kernel과 sample별 2D RoPE를 사용해 고정 token budget 아래 arbitrary-resolution latent grid와 variable-length text를 평탄화하며, packed classifier-free guidance는 1.09×–1.15× inference speedup을 보고한다.
Conditional 및 unconditional branch를 packing하면 명시된 denoising trajectory를 유지하면서 보고된 inference time이 감소하며, 이득은 1.09×에서 1.15×다.
3.2.2 Conditioning and rectified-flow training
고정된 Qwen3-VL-4B-Instruct가 text embedding을 제공하고, Mage-VAE는 이미지를 rectified-flow 학습용 16× reduced 128-channel latent로 변환한다. 편집에서는 같은 backbone이 instruction embedding, source-image latent, noisy target latent를 처리하며, frame-aware 3D RoPE가 source와 target token을 분리하고 loss는 target에만 계산한다.
3.3 Training Infrastructure
구현은 Mage-VAE, Qwen3-VL, NR-MMDiT에서 normalization, modulation, RoPE, gating, activation, residual 연산을 포함한 memory-bound 연산 체인을 융합한다. 8-GPU NVIDIA B200 node 1대에서 전체 구성은 MFU를 13.88%에서 29.28%로 높이고 GPU당 peak memory를 175.45 GB에서 141.44 GB로 낮추며 step time을 1.9285 s에서 0.7775 s로 줄여, 보고된 2.48× speedup을 달성한다.
Ablation은 tokenizer와 fusion 효과를 분리한다. Mage-VAE만으로 1.41× speedup을 얻고, 완전한 fused 구성은 더 낮은 memory로 2.48×에 도달한다.
4 Data Collection and Curation
데이터 전략은 text-to-image pair와 source-instruction-target 편집 triple을 위한 상호보완적 pipeline을 사용한다. 두 pipeline 모두 raw web data에서 충분히 나타나지 않는 시각 품질, 안전성, 의미 정렬, 다양성, capability coverage를 목표로 한다.
4.1 Text-to-Image Data Collection and Filtering
약 10B open-source image-text pair에서 시작한 생성 pipeline은 file 및 content filtering, SSCD/FAISS 기반 deduplication, Qwen3-VL-32B-Instruct multi-granularity recaptioning, concept-aware synthesis를 적용해 약 1.3B pair를 유지한다. 이후 단계에서는 resolution, watermark, aesthetic, OCR, safety threshold를 강화하고 long-text rendering, rare concept, layout, style 데이터를 보완한다.
Pipeline은 filtering 및 deduplication부터 multi-level recaptioning과 concept-aware synthesis까지 1.3B-pair corpus의 흐름을 추적한다.
나열된 threshold는 이후 단계에서 더 엄격해지며, 특히 resolution, watermark probability, aesthetic score에서 그렇다.
Sunburst는 생성 데이터가 Object & Products와 Scene & Place에 집중되는 반면, 편집 데이터는 semantic, attribute, text, object, style, low-level operation에 걸쳐 분포함을 보여준다.
4.2 Image-Editing Data Collection and Filtering
편집 pool은 약 90M triple로 시작하며, 이 중 약 50M은 open-source dataset에서, 40M은 내부적으로 합성했다. Qwen3.5-9B expert 3개가 instruction execution, preservation, plausibility를 판정하고, majority voting으로 약 45M triple을 유지한 뒤 19개 edit category로 매핑하고 rebalancing한다.
도식은 VLM expert 3개가 majority vote로 editing triple을 필터링한 뒤 taxonomy assignment와 category balancing을 수행함을 보여준다.
5 Training
생성과 편집은 Mage-VAE latent space, 4B NR-MMDiT, rectified-flow 학습을 공유한다. 학습 tree는 먼저 Mage-Flow-Base를 만들고, 이어서 aligned 및 Turbo 생성 모델과 source-conditioned Base, aligned, Turbo 편집 모델로 분기한다.
학습 tree는 progressive generation training, editing continuation training, Diffusion-NFT post-training, four-step distillation을 구분한다.
5.1 Pre-training and Supervised Fine-tuning
생성 학습은 fixed 256 × 256에서 1.2B pair, 512-pixel native-aspect-ratio regime에서 600M 고품질 sample, 1024-pixel regime에서 300M 더 정제된 sample, 150M curated pair의 SFT 순으로 진행한다. 편집은 Mage-Flow-Base에서 초기화한 뒤, 35M editing triple과 35M generation pair, 이어서 20M editing triple과 10M generation pair로 학습해 적응 과정에서도 open-ended generation prior를 보존한다.
이 visual은 semantic edit, restoration, stylization, camera change, structural output을 생성하는 통합 image-and-text-conditioned editor를 제시한다.
5.2 Diffusion-NFT Post-training
Diffusion-NFT는 likelihood estimation 없이 forward process에서 negative-aware online fine-tuning을 수행한다. 생성은 약 20K prompt를 OCR, aesthetic, semantic evaluator로 routing하며, 편집은 4:1 editing-to-generation update ratio, 약 30K editing prompt, instruction adherence, preservation, plausibility, text rendering을 위한 RationalRewards를 사용한다.
5.3 Few-step Distillation
Turbo 변형은 frozen DINOv2 및 CLIP feature에 기반한 Decoupled DMD와 adversarial perceptual guidance로 aligned teacher를 four-step student로 distill한다. 보고된 ablation에서는 대부분의 생성 및 text-editing metric이 개선되지만 general-editing 변화는 혼재하며, generation data를 추가하면 GEdit가 일관되게 개선되지는 않는 가운데 Turbo ImgEdit는 4.20에서 4.38로 증가한다. CLIP에 대해서는 이 글을 참조하라.
Schematic은 Decoupled DMD의 classifier-free augmentation 및 distribution-matching signal을 feature-space adversarial guidance와 분리해 보여준다.
Adversarial guidance는 보고된 많은 생성 및 text-editing score를 개선하지만, 보고된 general-editing 변화는 benchmark에 따라 다르다.
Generation-data mixing은 Turbo ImgEdit를 4.20에서 4.38로 높이지만 English 및 Chinese GEdit 변화는 균일하지 않다.
두 sunburst는 people, scene, product, text, control, complex edit를 포함한 생성 및 편집 distillation data의 category coverage를 보여준다.
6 Experiments
실험은 quantitative benchmark와 qualitative gallery를 통해 생성 및 편집 Base, aligned, Turbo 변형을 비교한다. 평가는 단일 benchmark가 아니라 품질, prompt following, text rendering, 편집 정확도, latency, memory를 다룬다.
6.1 Experimental Setup
생성은 native 1024²에서 GenEval, DPG-Bench, TIIF-Bench, OneIG, CVTG-2K, LongText로 평가하며, Base, aligned, Turbo 변형은 각각 30, 20, 4 step을 사용한다. 편집은 ImgEdit-Bench, GEdit-Bench, TextEdit-Bench로 평가하며, Base와 aligned 모델은 30 step, Turbo는 4 step을 사용한다.
6.2 Quantitative Results
Mage-Flow는 20 step에서 0.90 GenEval score, four steps에서 0.88을 보고하며, CVTG-2K 평균은 각각 0.887과 0.873이다. Mage-Flow-Edit-Turbo는 ImgEdit에서 4.38, GEdit-Bench English/Chinese에서 8.271/8.264, TextEdit-Bench Synthetic/Real-World에서 12.77/15.41을 보고하며, 논문은 정밀한 layout preservation과 복잡한 text replacement를 남은 한계로 지목한다.
Mage-Flow는 20 step에서 0.90 GenEval을 보고하고 Turbo는 four steps에서 0.88을 보고하며, 상대 성능은 8개 benchmark에 따라 달라진다.
세부 결과는 single-object, two-object, counting, positional GenEval category에서 강한 Mage-Flow score를 보고하며, DPG-Bench에서는 일관된 선두가 아닌 경쟁력 있는 성능을 보인다.
TIIF-Bench는 basic, advanced, designer task에서 short 및 long prompt를 구분하며, task category와 step budget에 따라 다른 동작을 보여준다.
Mage-Flow는 0.887 CVTG-2K 평균을, Turbo는 0.873을 보고하며, score는 2개부터 5개까지 text region별로 나뉜다.
OneIG는 Mage-Flow 변형의 강한 alignment 및 text subscore를 보고하지만, 일부 경쟁 모델보다 낮은 diversity와 낮은 Chinese overall score도 보인다.
요약 표는 보고된 step count에서 편집 시스템을 비교하며, 4B Mage 변형과 더 큰 모델 사이의 quality-efficiency trade-off를 보여준다.
Category-level ImgEdit 결과는 Turbo가 overall 4.38을 기록하며 extraction, hybrid, style, 기타 edit class에서 큰 차이를 보임을 보여준다.
Mage-Flow-Edit-Turbo는 semantic-consistency 및 perceptual-quality component와 함께 English에서 8.271, Chinese에서 8.264의 GEdit overall score를 보고한다.
Mage 변형 중 aligned 30-step editor가 TextEdit-Bench에서 가장 강하며, 14.14 Synthetic 및 16.26 Real-World를 보고한다. Turbo는 12.77과 15.41을 기록한다.
Gallery는 close-up 및 wide scene 전반에서 crop하지 않은 native-aspect-ratio 음식 및 landmark 생성을 보여준다.
Example은 wildlife, scenery, sculpture, surreal composite를 다루며 subject와 style의 폭을 보여준다.
Portrait gallery는 다양한 setting, age, cultural dress, lighting condition에서 사람을 보여준다.
Example은 text-rendering 평가를 보완하는 packaging, poster, cover, sign의 multi-line English text에 초점을 맞춘다.
Chinese 및 bilingual signage, packaging, cover layout은 multilingual text rendering의 qualitative example을 제공한다.
Paired source-output example은 localized text replacement, object counting, addition, removal, replacement, extraction을 보여준다.
Panel은 background replacement, zoom, viewpoint, scale, action, composite scene change를 보여준다.
Example은 landscape, food, architecture, object의 tone, style, material, color transformation을 보여준다.
Gallery는 virtual try-on, portrait retouching, hairstyle edit, meme composition, multi-instruction transformation을 포함한다.
Paired example은 bidirectional lens flare, rain, haze, blur, grayscale, colorization, exposure, enhancement task를 다룬다.
Example은 sketch, pose, Canny 및 HED map, depth, segmentation, normal map, control 기반 image reconstruction을 다룬다.
부록
- Appendix A는 Mage-VAE objective와 joint optimization을 상세히 설명하며, Appendix B는 80GB A100에서 4096² encoding/decoding이 149.6/375.0 ms라고 보고한다. Appendix는 reward evaluator도 명시하고 scientific-diagram fine-tuning을 설명하며, Mage-Flow-SciForma는 SciFormaBench-2K에서 zero-shot Mage-Flow-Base의 40.80 대비 61.61을 보고한다.
짧은 생각
논문의 효율성 주장은 backbone만 확장하기보다 tokenizer, native-resolution packing, fused kernel을 co-design하는 데 기반한다. 비교와 여러 evaluation benchmark는 보고된 점수 또는 model-judged score에 의존하며, 저자들은 불완전한 multilingual long-text rendering, 제한적인 multi-image editing coverage, general editing에서 균일하지 않은 distillation gain을 언급한다.