Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis | Summary
31 Mar 2026 | Paper Review Text-to-Image Generation Search Agents Supervised Fine-TuningContents
- Summary
- 1 Introduction
- 2 Related Works
- 3 Preliminary
- 4 Data Pipeline
- 5 Methodology
- 6 Experiments
- 7 Conclusion
- Limitations and Future Work
- Appendix
- Brief Thoughts
This article explains the key points of Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis.
- 2026-03-31 (arXiv)
- Chen, Shuang, Shou, Quanxin, Chen, Hangting, Zhou, Yucheng, Feng, Kaituo, Hu, Wenbo, Zhang, Yi-Fan, Lin, Yunlong, Huang, Wenxuan, Song, Mingyang, et al.
- University of California, Los Angeles, Tencent Hunyuan, The Chinese University of Hong Kong, The Hong Kong University of Science and Technology
- Paper
- Github
Summary
- Unify-Agent addresses factual image generation for rare identities and long-tail concepts whose defining appearance is poorly represented in model parameters. It adapts Bagel into a Think–Research–Recaption–Generate pipeline that retrieves textual and visual evidence, separates identity requirements from scene requirements, and synthesizes images from a grounded specification and selected references.
- The authors construct 143K verified trajectory-image pairs and introduce FactIP, containing 2,462 prompts across 12 subcategories; reported FactIP results use the 500-prompt FactIP-Mini split. Unify-Agent raises FactIP Overall from Bagel’s 50.9 to 73.2 and reports 0.77 on WiSE, 4.08 on KiTTEN, and 77.4 and 71.5 on the SKCI and MKCC levels of T2I-FactBench.
- Removing image search reduces FactIP Overall to 56.2, while removing recaptioning reduces it to 62.9, supporting the value of visual evidence and structured evidence integration. The system remains a relatively shallow workflow, evaluation relies on automated judges, and the strongest commercial models in the comparisons still lead on FactIP and WiSE.
1 Introduction
The paper distinguishes visual plausibility from factual fidelity: a convincing image can still depict the wrong person, character, product, or cultural object. Its proposed remedy is inference-time access to external knowledge followed by a learned conversion of that knowledge into visual constraints.
- Unify-Agent places prompt interpretation, evidence-seeking actions, recaptioning, and generation within a shared unified multimodal model, while retaining external search tools.
- The recaption separates identity-defining features from user-controlled pose, clothing, environment, composition, and style.
- Figure 1 illustrates the range of generated subjects and styles; selected examples alone do not establish general reliability.
- The project URL is https://github.com/shawn0728/Unify-Agent.
The montage spans landmarks, real people, fictional characters, food, and designer toys, illustrating the intended breadth of identity-sensitive synthesis and stylistic adaptation. Selected outputs without controlled comparisons demonstrate examples, not an estimated success rate.
2 Related Works
The related-work discussion connects unified understanding-generation architectures, agentic workflows with external tools, and factual evaluation of image synthesis. Unify-Agent integrates these strands through supervised evidence acquisition and grounded recaptioning.
2.1 Unified multimodal models
Unified multimodal models such as the Janus family and Show-o share architectural resources between visual understanding and synthesis. The authors argue that these capabilities do not by themselves resolve long-tail factual generation when inference depends on knowledge internalized during pretraining.
- Hallucinated appearance and identity drift motivate acquiring external evidence before synthesis.
- The relevant distinction is inference-time grounding, not simply whether understanding and generation share a backbone.
2.2 Agentic workflows
Prior agentic image-generation systems often connect a planner, retrieval tools, and a separate generator through multiple APIs. The authors identify weak evidence integration and cascading errors as limitations and instead train a unified model to produce the research and recaptioning trajectory.
- Unification applies to the learned reasoning and generation components; external text and image retrieval remain part of execution.
- The experiments do not include a separately trained planner-generator pipeline matched for training data and retrieval access, so they do not isolate architectural unification from these other advantages.
2.3 World knowledge and factual evaluation.
T2I-FactualBench, KiTTEN, and WiSE evaluate knowledge-sensitive generation beyond aesthetics and generic prompt alignment. FactIP adds reference-based assessment of rare identities and culturally significant concepts, emphasizing preservation of defining visual attributes.
3 Preliminary
The preliminary section establishes Bagel’s text and image objectives, tests external knowledge injection without additional training, and formulates synthesis as a trajectory with explicit intermediate grounding variables. These components motivate the evidence-acquisition policy and recaption interface.
3.1 Unified Multimodal Model: Bagel
Bagel uses a Mixture-of-Transformers architecture with dedicated understanding and generation experts. Multimodal Understanding uses autoregressive next-token prediction, whereas Multimodal Generation predicts a rectified-flow velocity field in a continuous VAE latent space.
- The text objective minimizes the negative log-likelihood of target tokens conditioned on preceding tokens and multimodal context C.
- The image objective minimizes expected squared error between predicted and target latent velocity fields, with the continuous timestep sampled uniformly from (0, 1).
- Unify-Agent retains both objectives during adaptation to agent trajectories.
3.2 Motivating Evidence for World-Grounded Synthesis
A training-free study uses 200 FactIP examples from Scene, Character, and Object to compare prompt-only Bagel with text injection, two ground-truth visual references, and joint text-plus-visual injection. Figure 2 reports overall absolute gains of 3.9, 15.0, and 14.0, respectively.
- Visual injection improves Character by 16.6 and Object by 16.3, compared with 3.1 and 7.9 for text injection.
- Joint injection underperforms visual-only injection overall, motivating evidence organization rather than assuming that additional raw conditioning always helps.
- Because the visual intervention supplies ground-truth references, it measures the potential benefit of appearance evidence rather than the reliability of autonomous retrieval.
In the 200-example training-free study, two ground-truth visual references produce a 15.0-point overall gain over prompt-only Bagel, compared with 3.9 for text and 14.0 for combined injection. The weaker joint gain motivates evidence organization, but the experiment does not test autonomous retrieval quality.
3.3 Problem Formulation
The formulation introduces cognitive gap assessment g, textual evidence trace τt, visual evidence trace τv, and grounded recaption c before producing image y. Equation 3 factorizes the trajectory as pθ(y, c, τt, τv, g | x) = pθ(g | x) · pθ(τt, τv | x, g) · pθ(c | x, g, τt, τv) · pθ(y | c, τv).
- The factors represent gap detection, evidence acquisition, recaptioning, and visual synthesis.
- The final generation factor excludes direct conditioning on the original prompt and upstream textual reasoning history; relevant information must reach generation through the recaption and visual evidence.
- This conditional-independence design is a modeling assumption implemented through attention restrictions, not an independently established property of factual generation.
4 Data Pipeline
The data pipeline produces process supervision for agent fine-tuning and a separate evaluation benchmark. Figure 3 traces concept collection, prompt and reference preparation, multimodal research, recaption annotation, and generation-based verification.
- Training supervision covers evidence acquisition and transformation as well as the final synthesis target.
- The benchmark comes from the same task pool but uses a disjoint subset excluded from supervised fine-tuning.
4.1 Training Data Construction
Training Data Construction begins with 456K manually verified examples across 12 domains, using BabelNet and Wikipedia information, two representative seed images, GPT-4o metadata summaries, and Gemini 3 Pro prompts. Research construction and generation-based verification retain 143K trajectory-image pairs.
- 4.1.1 Task Source and Prompt Collection: annotators remove incorrect metadata, poor references, and overly common concepts; generated prompts vary pose, clothing, scene, lighting, and style while requiring identity preservation.
- 4.1.2 Multimodal Research Trace Construction: Claude Opus 4.6 constructs textual research followed by visual research, allowing semantic disambiguation to guide image queries.
- Textual research records the issued query and retrieved evidence summary. Visual research conditions its query on the prompt, gap assessment, and preceding textual grounding.
- The prose identifies Gemini 3 Flash as the candidate-image evaluator, scoring identity consistency, subject salience, clarity, and watermark cleanliness before selecting the top two images; Figure 3 instead labels the retrieval judge Claude Opus 4.6.
- Trace representation and supervision records textual query and evidence followed by visual query and selected images, making evidence acquisition an explicit supervision target.
- 4.1.3 Evidence-Grounded Recaption Annotation combines user composition, semantic background, and identity-critical visual cues into a generation specification.
- Nano Banana Pro generates validation images from the recaption and two references; GPT-4o checks identity consistency against seed imagery. Rejection sampling permits up to five trials, after which an unsuccessful trajectory is discarded.
The diagram shows supervision for research and evidence transformation before generation, followed by generation-based rejection sampling. It separates benchmark preparation from training-data production; its Claude Opus 4.6 retrieval-judge label differs from the Gemini 3 Flash evaluator named in the prose.
4.2 FactIP Evaluation Benchmark
FactIP evaluates long-tail, visually identifiable concepts across Character, Object, and Scene, with 12 fine-grained subcategories. The full benchmark contains 2,462 prompts, and FactIP-Mini samples 500 while preserving the category distribution; Appendix D.2 states that all reported FactIP results use this mini split.
- Each sample contains a user instruction, two ground-truth seed images, and target metadata; evaluation instances are excluded from fine-tuning.
- Selection prioritizes knowledge-intensive concepts, clear visual identities, difficulty for memorization-only generation, and category diversity.
- Seed2.0 independently scores Clarity, Content, Aesthetics, and Relevance on an integer 0–10 scale while considering both references jointly.
- The weighted score assigns 0.05 to Clarity, 0.10 to Content, 0.10 to Aesthetics, and 0.75 to Relevance, then multiplies the result by 10 to obtain a [0, 100] score.
- Section 4.2.1 describes 2,500 manually filtered samples; the introduction and Appendix D identify 2,462 as the full benchmark size without detailing the reduction.
5 Methodology
The methodology combines supervised learning on interleaved agent trajectories with inference-time evidence acquisition and integration. Its central interface is a grounded recaption that separates identity constraints from the user’s requested composition and style.
5.1 Unified Fine-Tuning on Multimodal Agent Trajectories
Unified Fine-Tuning on Multimodal Agent Trajectories packs heterogeneous samples into sequences and optimizes LSFT = Ltext + Limage. Text supervision covers reasoning, tool invocation, and recaptioning, while image supervision retains latent flow matching with the VAE frozen.
- Special trajectory tokens such as <think>, <tool_call>, and <recaption> receive increased cross-entropy weight to reinforce structural boundaries and executable formats.
- Separate supervision masks control the objectives: text-only trajectory segments contribute to Ltext, whereas image-generation samples contribute to both Ltext and Limage.
- The hybrid mask applies causal attention to textual reasoning and full attention within reference visual-token blocks.
- Generation tokens retain access to reference visual tokens and recaption tokens but not direct access to historical research text or the original textual input. Figure 6 visualizes this restriction, which does not exclude indirect transmission of upstream information through contextualized representations.
5.2 Unify-Agent: Inference-Time Multimodal Agentic Pipeline
At inference, Think decomposes the prompt and identifies missing visually consequential information; Research acquires textual grounding before searching for compatible visual references. Recaption converts the evidence into identity-preserving and scene-compositional constraints, and Generate synthesizes from the recaption and visual anchors.
- Think represents missing information as knowledge units such as facial structure, hairstyle, signature appearance, object geometry, or factual scene details; the model may bypass research when no gap is identified.
- Research conditions visual queries on prompt intent and preceding textual evidence rather than only an entity name.
- Recaption retains target-specific appearance while specifying pose, environment, garment, mood, composition, and presentation.
- Generate blocks direct conditioning on the full upstream textual history, relying on the recaption for controllable description and visual anchors for appearance.
- Figure 4 illustrates these stages with a DUDOO popcorn-bucket figure and a macro-photography request.
The DUDOO example traces conversion of a short product prompt into a detailed synthesis specification. Appearance attributes extracted from selected references are separated from macro-photography and scene requirements before integration into the recaption.
6 Experiments
Experiments evaluate reference-grounded identity synthesis, broad world knowledge, fine-grained entity alignment, and factual concept composition. Benchmark comparisons are accompanied by pipeline, constraint, and recaption-stage visual-representation ablations.
6.1 Experimental Details
The comparisons include commercial generators, generation-only models, and unified multimodal models, with Bagel serving as the base-model reference. GPT-4o judges WiSE, KiTTEN, and T2I-FactualBench, whereas Seed2.0 judges FactIP.
- Commercial references include the Seedream family, Nano Banana variants, DALLE-3, and GPT-Image models; other comparisons include FLUX, Stable Diffusion, Qwen-Image, Janus, Emu, and Hunyuan-Image.
- Appendix A reports Bagel-14B fine-tuning on 64 NVIDIA H20 GPUs for approximately 10 days and 10,000 gradient steps.
- Training uses AdamW, a constant learning rate of 5 × 10−5 after 500 warmup steps, maximum gradient norm 5.0, equal CE and MSE weights of 1.0, and special-token CE weight 3.0.
- The paper does not report matched retrieval budgets, inference-latency comparisons, confidence intervals, or an aggregate human validation study of judge agreement.
6.2 Main Results
Tables 1–4 report improvements over Bagel across all four benchmarks, with FactIP gains concentrated in identity relevance. Unify-Agent leads the listed unified models on the headline scores, but does not outperform the strongest commercial generators across these evaluations.
- FactIP: Overall rises from 50.9 to 73.2, a 22.3-point improvement. Character, Object, and Scene Relevance reach 67.3, 71.8, and 78.2; Nano Banana-2 and Seedream-5 remain higher overall at 88.5 and 87.3.
- WiSE: Unify-Agent reaches 0.77 versus BAGEL’s 0.52 and BAGEL+CoT’s 0.70. Cultural, Biology, and Chemistry scores are 0.82, 0.72, and 0.70; BAGEL+CoT is slightly higher on Space and Physics, while Nano Banana reaches 0.89 overall.
- KiTTEN: Text alignment is 4.22, Entity alignment 3.93, and Overall 4.08, compared with Imagen-3’s 4.17, 2.83, and 3.50. Unify-Agent leads the listed models overall, although individual category scores do not all lead.
- T2I-FactBench: SKCM Concept reaches 69.2, SKCI All 77.4, and MKCC All 71.5, compared with Bagel’s 31.8, 58.2, and 66.3. DALLE-3 remains higher on SKCI All at 80.5 and MKCC All at 75.7.
- The KiTTEN Bagel-CoT row reports aggregate Entity alignment of 2.02, below every category-level entity score in that row; this aggregate cannot be reconciled with a nonnegative weighted average of the displayed categories.
Unify-Agent has the highest Overall score among the listed unified models at 73.2, compared with Bagel’s 50.9. Category-level Relevance improves substantially, while Nano Banana-2 and Seedream-5 remain higher overall at 88.5 and 87.3.
Unify-Agent improves WiSE Overall to 0.77 from BAGEL’s 0.52 and BAGEL+CoT’s 0.70. It leads the unified group in Cultural, Time, Biology, and Chemistry, but not Space or Physics, and remains below the listed commercial models overall.
Unify-Agent reports Text alignment 4.22, Entity alignment 3.93, and Overall 4.08, leading the displayed aggregates without leading every category. The Bagel-CoT aggregate Entity score of 2.02 is below all its displayed category-level Entity scores, indicating an unresolved reporting inconsistency.
SKCM Concept rises from Bagel’s 31.8 to 69.2, while SKCI All and MKCC All reach 77.4 and 71.5. Multi-concept composition remains uneven: Unify-Agent’s Composition score of 73.6 is below Bagel-CoT’s 77.1 and DALLE-3’s 85.6.
6.3 Analysis of Experimental Results
Table 5 shows that the full pipeline improves FactIP Relevance from 44.9 to 72.4 while leaving Clarity nearly unchanged at 91.2 versus vanilla Bagel’s 91.3. The measured benefit therefore concerns factual identity fidelity more than uniformly better perceptual quality.
- Pipeline ablations: removing Text-Search, Image-Search, or Recaption reduces Overall from 73.2 to 65.4, 56.2, or 62.9, respectively.
- Removing Image-Search produces the largest pipeline-level relevance loss, from 72.4 to 50.8, consistent with the importance of direct appearance anchors.
- Removing Recaption also lowers Clarity to 83.0 and Aesthetics to 74.5, suggesting that structured evidence integration helps synthesis quality as well as identity matching.
- Constraint ablations: removing identity-preserving constraints lowers Relevance to 65.9; removing scene-compositional constraints lowers Content to 70.8 and Aesthetics to 80.7.
- The surrounding prose claims that the full model is best on all metrics, but the table gives the no-image-search variant higher Clarity at 92.1.
Removing image search causes the largest pipeline-level Overall loss, while removing recaptioning also reduces clarity and aesthetics. The ViT ablation is more damaging than the VAE ablation, supporting complementary visual representations; these generation scores do not directly establish improved standalone understanding.
6.4 Finding: Why Generation Helps Understanding in a Unified Model
The authors interpret recaption architecture ablations as evidence that generative representations can assist multimodal understanding. Removing VAE inputs reduces Overall from 73.2 to 71.2, while removing ViT inputs reduces it to 61.4, indicating complementary but unequal contributions.
- ViT removal lowers Relevance to 58.7, compared with 70.8 after VAE removal, making the semantic pathway more consequential in this experiment.
- The proposed explanation is that ViT tokens support identity and relational interpretation, while VAE latents retain texture, material, color, and local structure useful for recaptioning.
- These downstream generation scores support the utility of dual visual representations within the system; they do not directly measure standalone understanding or establish a general benefit from generation training.
- 6.4.1 Visualization comparison results: Figure 5 compares historical figures, LIVE A LIVE’s Steel Titan, a Zaku Tank model, DUDOO, and Jacques Derrida. The selected outputs illustrate retention of subject-specific appearance relative to generic or mismatched baseline outputs.
The selected comparisons contrast target-specific faces, fictional robot designs, collectible-product structures, and toy appearance with generic or mismatched baseline outputs. They illustrate the intended identity-grounding benefit, but the selection and success markers are not an independent quantitative evaluation.
7 Conclusion
The conclusion frames factual synthesis as a multimodal reasoning and evidence-integration task. Its empirical basis is improvement over Bagel across four benchmarks and performance losses when retrieval, recaptioning, or grounded constraints are removed.
Limitations and Future Work
The authors acknowledge Bagel’s limited long-context capacity and support for relatively few images, which constrain complex agent behavior. They characterize the method as a relatively shallow one-pass workflow rather than an iterative agent with sustained interleaved search, reflection, and replanning.
- The demonstrated scope is primarily IP- and concept-centric synthesis, not longer-horizon tasks such as travel planning or academic report generation.
- Proposed extensions include stronger unified backbones, repeated search, longer-horizon planning, and adaptive reasoning.
- Commercial-model gaps and reliance on automated evaluation further limit broad interpretations of the benchmark improvements.
Appendix
- A Implementation Details: Table 6 specifies Bagel-14B, SigLIP-SO400M-14 (NaViT), a frozen FLUX VAE, ViT patch size 14, and VAE latent patch size 2. Maximum tokens per sample and expected tokens per batch are both 40,240, with a hard packed-token limit of 41,520; distributed training uses FSDP (HYBRID_SHARD). Fine-tuning updates the language model, ViT encoder, and multimodal connectors while freezing the VAE.
- B Attention Masking Strategy: Figure 6 shows causal attention for textual reasoning, full attention within reference visual-token blocks, and restricted generation conditioning on reference images and recaptions. The mask blocks direct access to upstream textual history, but useful or noisy upstream information can still propagate through the recaption or contextualized reference representations.
- C Evaluation Protocols: WiSE uses GPT-4o ratings for Consistency, Realism, and Aesthetic Quality on a 0–2 scale, weighted 0.7, 0.2, and 0.1; the appendix does not explicitly explain normalization to the scale reported in Table 2. KiTTEN independently scores Text Alignment and Entity Alignment from 1 to 5, although Table 3’s caption states 0 to 5 and adds an aggregate score. T2I-FactBench uses successive binary VQA checks for concept shape, color, texture, feature details, instantiation, and, for MKCC, composition.
- FactIP Benchmark Evaluation assigns Seed2.0 four independent 0–10 ratings, places 0.75 of the weighted score on Relevance, and multiplies the weighted result by 10. Stable identity cues are inferred jointly from GT1 and GT2, without over-penalizing variations already present across the two references.
- D More Details about FactIP benchmark: Table 7 lists Animation 438, Comic 363, Celebrity 300, Game 272, Mascot 77, Mythology 50, Food 316, Cultural Relic / Art 126, Toy 123, Animal / Plant 50, Landmark 297, and Festival / Celebration 50 prompts, totaling 2,462. Figure 7 presents the hierarchy and category-wise comparisons, although its displayed group counts do not match Table 7. D.2 states that all results use FactIP-Mini; Table 8 reports Unify-Agent Overall as 71.7 rather than the 73.2 in Tables 1 and 5, without explaining the discrepancy.
- E Prompt documents E.1 System Prompt, E.2 Evluation Prompt, E.3 Judge Prompt, E.4 Prompt-Generation Prompt, E.5 Recaption Prompt, and E.6 Summary Prompt. The templates describe text_search and search_image interfaces, visual-reference filtering, nationality-based prompt-language selection, explicit image_1 and image_2 grounding, preservation statements, and query-focused webpage summaries. E.2 requests five scores and six fields, but lists four evaluation dimensions and displays four scores plus a rationale.
- F Image Generation Showcases presents Copper Combustion, DUDOO Tasty Afternoon Tea Series, William Butler Yeats, Grigory Perelman, and Bruce Beutler in Figures 8–12. The traces respectively illustrate green-flame evidence retrieval, toy-identity and product-detail integration, portrait grounding for a cottage-writing scene, appearance extraction for an apartment scene, and reference-guided adaptation to a laboratory setting. These are demonstrations rather than independent factual verification: the copper example mixes combustion and copper-compound flame-test descriptions, and the DUDOO example’s second search returns no relevant pages while its displayed search_image call conflicts with the text-search stage label.
- G FactIP Benchmark Evaluation Showcases presents Gregg Popovich, Max Weber, Habatan, and Scottie Pippen in Figures 13–16, with displayed Overall ratings of 9, 5, 9, and 2. The evaluator rewards Popovich’s likeness and basketball context, penalizes Weber’s identity mismatch and modern microphone, praises Habatan’s stylistic rendering while alleging a shoe-color discrepancy, and penalizes Pippen’s likeness, artifacts, and jersey details. The displayed Overall ratings do not directly follow Appendix C’s weighted formula, and the Habatan shoe-color allegation is not supported by the shown references.
Brief Thoughts
The clearest contribution is the recaption interface: retrieved evidence becomes explicit identity and composition constraints before generation, and the no-recaption ablation supports this design. The broader claim that generation improves understanding is less directly established because the supporting experiment measures downstream synthesis rather than understanding in isolation. A decoupled agent matched for training data, references, and retrieval budgets would help isolate the value of architectural unification. Judge calibration, consistent aggregate reporting, and an explanation of the FactIP 73.2 versus 71.7 discrepancy are important for reproducibility. The results establish a substantial improvement over Bagel on knowledge-intensive synthesis, not a general solution to reliable open-world multimodal agency.