Gorio Tech Blog search

Image Generators are Generalist Vision Learners | Summary

|

Contents

This article explains the key points of Image Generators are Generalist Vision Learners.

  • 2026-04-22 (arXiv)
  • Gabeur, Valentin, Long, Shangbang, Peng, Songyou, Voigtlaender, Paul, Sun, Shuyang, Bao, Yanan, Truong, Karen, Wang, Zhicheng, Zhou, Wenlei, Barron, Jonathan T., et al.
  • Google DeepMind
  • Paper

Read this article in Korean


Summary

  • Image Generators are Generalist Vision Learners investigates whether image-generation pretraining supports competitive visual understanding without task-specific architectures. Vision Banana instruction-tunes Nano Banana Pro on its original generation mixture plus a low proportion of vision-task data, representing segmentation, metric depth, and surface normals as decodable RGB images.
  • Under zero-shot transfer to evaluation benchmarks, Vision Banana achieves 69.9 mIoU on Cityscapes, 73.8 cIoU on RefCOCOg UMD val, and 79.3 gIoU on ReasonSeg val. The ReasonSeg result uses Gemini 2.5 Pro to interpret queries, while the SA-Co/Gold result of 47.5 cgF1 uses Gemini 3.1 Flash-Lite for object-presence filtering.
  • Metric depth instruction-tuning uses synthetic depth supervision and no camera parameters; depth inference also requires no camera parameters. Vision Banana achieves average δ1 of 0.882 across six datasets and 0.929 on the four datasets reported for Depth Anything 3, versus 0.918 for that specialist. Surface-normal prediction achieves the lowest reported average indoor mean angular error, 15.549, among the compared methods, but does not lead on every dataset.
  • Human preference win rates against Nano Banana Pro are 53.5% for text-to-image generation and 47.8% for image editing, consistent with retention of generation capabilities but not sufficient to establish statistical equivalence. Limitations include higher computational overhead than lightweight specialists, instance-merging and group-splitting failures, external-model assistance on two benchmarks, and unreported tuning quantities and costs.

1. Introduction

The paper tests whether a pretrained image generator can produce perception outputs precise enough for decoding and quantitative evaluation. Prior approaches either generated visualizations without reliable format compliance or adapted generators into task-specific systems, limiting their generality across perception and generation.

  • The paper adopts the language-model distinction between generative pretraining and instruction tuning: the latter aligns an existing generator to requested tasks and output formats.
  • Figure 1 illustrates shared RGB outputs for five perception tasks alongside text-to-image generation and editing; Table 1 summarizes the quantitative evidence.
  • Project Page: vision-banana.github.io

The diagram shows how changing the instruction yields different perception visualizations without replacing the generator. Its radar comparison summarizes selected results; the accompanying tables are needed to distinguish training settings and auxiliary-model use.

Unified generation and visual-understanding tasks supported by Vision Banana through prompted RGB outputs.
Unified generation and visual-understanding tasks supported by Vision Banana through prompted RGB outputs.

The overview combines segmentation, geometry, and generation comparisons. Its depth value is a four-dataset average rather than the six-dataset average in the detailed results, and its instance and ReasonSeg entries summarize assisted pipelines.

Headline benchmark comparisons for 2D understanding, 3D understanding, text-to-image generation, and image editing.
Headline benchmark comparisons for 2D understanding, 3D understanding, text-to-image generation, and image editing.

2. Method

Vision Banana mixes vision-task data into Nano Banana Pro’s original training mixture at a very low ratio. The model retains a common RGB output interface rather than introducing specialized architectures or custom task losses.

  • For 2D tasks, supervision comes from in-house model annotations of web-crawled images; 3D supervision comes from rendering engines.
  • The authors state that training data from the evaluation benchmarks is excluded from the instruction-tuning mixture. This does not establish the contents of the base model’s pretraining data.
  • The PDF does not report an exact mixing ratio, vision-example count, tuning duration, or compute budget, limiting independent assessment of the claimed lightweight adaptation.

Retaining image-generation data is intended to preserve the base model’s generative capabilities during perception alignment. Human comparisons yield 53.5% wins on GenAI-Bench and 47.8% on ImgEdit against Nano Banana Pro.

  • These results are consistent with broadly comparable generation and editing performance, rather than an improvement on both tasks.
  • The PDF does not provide preference-study sample sizes or uncertainty intervals.
  • Figures 11 and 12 supplement these aggregate results with paired generation and editing examples.

Both checkpoints generate detailed images for the same four prompts, supporting qualitative retention of synthesis capabilities. These selected pairs complement, but do not replace, the reported GenAI-Bench human preference result.

Paired Vision Banana and Nano Banana Pro text-to-image outputs for GenAI-Bench prompts.
Paired Vision Banana and Nano Banana Pro text-to-image outputs for GenAI-Bench prompts.

3. Vision Banana - Generalist Vision Model from Image Generator

The evaluation covers referring expression, semantic, and instance segmentation, together with monocular metric depth and surface-normal estimation. One instruction-tuned model supplies the perception outputs; prompts and decoding procedures change across tasks, and two benchmark pipelines additionally use a multimodal LLM.

  • “Zero-Shot Transfer” means that a method was not trained on the benchmark’s in-domain training split, not that Vision Banana received no task supervision.
  • The tables separate methods trained on benchmark data from transfer methods; this distinction is necessary when interpreting comparisons with specialists.

3.1. 2D Semantic Understanding

Semantic segmentation prompts specify category-to-color assignments, and each generated pixel is decoded to the nearest target color in RGB space. Cityscapes evaluation uses the same complete mapping for all 19 classes on every image, including classes absent from that image.

  • Table 2 reports 69.9 mIoU versus 65.2 for SAM 3, a 4.7-point improvement among the listed zero-shot transfer methods.
  • The benchmark-trained SegMan-L scores 84.2, so Vision Banana does not achieve the best result across all training settings.
  • Figure 2 illustrates named colors, RGB tuples, and hex codes, as well as queries for object parts and descriptive categories such as “patterns on the wall.”

Vision Banana leads the listed zero-shot transfer methods at 69.9 mIoU, exceeding SAM 3 by 4.7 points. SegMan-L’s 84.2 shows that this is not the overall best result when benchmark-specific training is allowed.

Cityscapes validation semantic-segmentation results separated by benchmark-training setting.
Cityscapes validation semantic-segmentation results separated by benchmark-training setting.

Different category and color specifications produce different maps from the same input images. The examples show support for object parts and descriptive phrases as well as conventional categories, but do not measure reliability across such queries.

Prompt-dependent semantic-segmentation examples with named colors, RGB tuples, and hex codes.
Prompt-dependent semantic-segmentation examples with named colors, RGB tuples, and hex codes.

Instance segmentation prompts specify a target class and background color while allowing the generator to choose a distinct color for each instance. Because the instance count and palette are unknown beforehand, a multi-stage clustering algorithm converts the generated visualization into discrete masks.

  • Figure 3 shows class-specific separation of garlic, beef, price tags, crescent-shaped croissants, basketballs, and balls.
  • SA-Co/Gold contains 168k Image-NP pairs, mostly negative queries. Since the model was not tuned to produce empty masks, Gemini 3.1 Flash-Lite first classifies target presence, and Vision Banana processes only predicted positives.
  • Table 3 reports 47.5 cgF1, 0.84 IL_MCC, and 56.0 pmF1 for the combined pipeline, versus 24.6, 0.57, and 42.0 for OWLv2. Benchmark-trained SAM 3 remains ahead at 54.1 cgF1, or 61.2 with fine-tuned Llama 3.2.

The Vision Banana + Gemini 3.1 Flash-Lite pipeline scores 47.5 cgF1, 0.84 IL_MCC, and 56.0 pmF1. Presence classification contributes to the gated score, while SA-Co-trained SAM 3 pipelines remain ahead overall.

SA-Co/Gold instance-segmentation comparison using gated, classification, and positive-mask metrics.
SA-Co/Gold instance-segmentation comparison using gated, classification, and positive-mask metrics.

The outputs distinguish instances using model-selected colors and change their targets when the noun phrase changes. The basketball-versus-ball examples illustrate that the requested class controls which instances appear.

Class-specific instance-segmentation visualizations with distinct colors for individual objects.
Class-specific instance-segmentation visualizations with distinct colors for individual objects.

Referring expression segmentation tests grounding of free-form object descriptions. Vision Banana scores 73.8 cIoU on RefCOCOg UMD val, slightly above the SAM 3 + Gemini 2.5 Pro pipeline at 73.4, while benchmark-trained X-SAM-Phi3-3.8B reaches 83.8.

  • On ReasonSeg val, Gemini 2.5 Pro translates the reasoning query into a descriptive reference before segmentation. Gemini and Vision Banana are each queried exactly once in the single-turn pipeline.
  • The assisted pipeline scores 79.3 gIoU versus 77.0 for SAM 3 + Gemini 2.5 Pro. Both pipelines use the same named reasoning assistant, but the paper does not establish that their intermediate references are identical.
  • Figure 4 illustrates grounding by appearance, action, atypical object use, and Chinese and English text. These selected examples demonstrate capability but do not quantify its frequency or reliability.

The 73.8 versus 73.4 cIoU comparison shows a small numerical advantage over SAM 3 + Gemini 2.5 Pro within zero-shot transfer. Higher benchmark-trained results delimit the scope of that advantage; uncertainty estimates are not reported.

RefCOCOg UMD validation referring-expression segmentation results.
RefCOCOg UMD validation referring-expression segmentation results.

With Gemini 2.5 Pro assisting both compared pipelines, Vision Banana reaches 79.3 gIoU versus 77.0 for SAM 3. This measures an assisted segmentation pipeline rather than standalone reasoning by Vision Banana.

ReasonSeg validation results for reasoning-based referring-expression segmentation.
ReasonSeg validation results for reasoning-based referring-expression segmentation.

The selected masks ground appearance, activity, unusual functional use, and multilingual text. The toaster example illustrates functional reference resolution, although the figure does not provide a systematic reasoning evaluation.

Referring-expression grounding examples involving appearance, actions, atypical object roles, and multilingual text.
Referring-expression grounding examples involving appearance, actions, atypical object roles, and multilingual text.

3.2. 3D Understanding from Monocular Images

Metric depth is encoded through the power transform of Barron (2025), followed by interpolation along a piecewise-linear path on the RGB cube’s edges. Equation (1) is (f(d,\lambda,c)=1-\left(1-\frac{d}{\lambda c}\right)^{\lambda+1}), with λ < −1; every experiment uses λ = −3 and c = 10/3.

  • The transform maps nonnegative metric distances into [0, 1), allocating more representational resolution to nearby depths.
  • Figure 5 shows the color path from black toward white. Generated colors are projected to the nearest path segment, after which the interpolation and depth transform are inverted.
  • The invertible correspondence is between depth and the selected color path, not all points in the RGB cube. Training also includes Plasma, Inferno, Viridis, and grayscale representations.

Distance labels along the RGB-cube path show the nonlinear allocation of color resolution, with closer depths receiving greater differentiation. Decoding projects generated colors onto this path before recovering metric distances; the mapping does not cover the entire RGB cube bijectively.

Power-transformed metric distances encoded along a piecewise-linear path on the RGB cube.
Power-transformed metric distances encoded along a piecewise-linear path on the RGB cube.

Table 6 evaluates metric depth on NYU, iBims1, ETH3D, DIODE-Indoor, KITTI, and nuScenes. Vision Banana’s depth-task supervision is entirely synthetic, and the authors report using no camera intrinsics or extrinsics during depth training or inference. It obtains average δ1 of 0.882 and AbsRel of 0.116.

  • Across the same four datasets reported for Depth Anything 3—NYU, ETH3D, DIODE, and KITTI—average δ1 is 0.929 versus 0.918.
  • Performance is uneven: Vision Banana scores 0.935 on ETH3D and 0.917 on DIODE-Indoor, but 0.643 on nuScenes, below UniK3D at 0.840 and MoGe-2 at 0.820.
  • Depth Anything 3 also leads Vision Banana on NYU and KITTI. The aggregate advantage does not imply uniform superiority across domains.

The table supports an aggregate depth advantage while exposing domain-specific weaknesses, notably δ1 of 0.643 on nuScenes. Its footnotes specify matched four-dataset comparisons and identify DepthLM’s nuScenes result as non-zero-shot.

Monocular metric-depth results across six benchmarks, including camera-intrinsics requirements and matched-subset averages.
Monocular metric-depth results across six benchmarks, including camera-intrinsics requirements and matched-subset averages.

Figure 6 decodes generated depth images and uses camera intrinsics to unproject them into point clouds, illustrating planar structure and object geometry in selected NYU v2 and ETH3D scenes. The intrinsics are used for reconstruction, not for predicting depth.

  • Figure 7 provides an informal smartphone-image check near Kinkaku-Ji: one marked point is predicted at 13.71 meters versus a Google Maps measurement of 12.87 meters, with reported AbsRel around 0.065.
  • This single-point check is illustrative evidence, not calibrated depth ground truth or a systematic evaluation of consumer-device generalization.

Novel views of unprojected depth predictions illustrate room surfaces and object placement in selected scenes. Camera intrinsics enter the point-cloud reconstruction step, not the depth-estimation model; these views are qualitative checks rather than independent reconstruction metrics.

Input images, generated metric-depth visualizations, and reconstructed point-cloud views from NYU v2 and ETH3D.
Input images, generated metric-depth visualizations, and reconstructed point-cloud views from NYU v2 and ETH3D.

The marked point compares a 13.71-meter prediction with a 12.87-meter Google Maps measurement. This offers an informal scale check, but only one location is measured and the reference is not a calibrated depth map.

Smartphone photograph near Kinkaku-Ji with predicted depth and a Google Maps distance check.
Smartphone photograph near Kinkaku-Ji with predicted depth and a Google Maps distance check.

Surface normals use camera-space coordinates with +x right, +y up, and +z pointing out of the image plane. Their components are encoded as R = trunc((1 − x)/2, min = 0, max = 1) × 255, G = trunc((1 + y)/2, min = 0, max = 1) × 255, and B = trunc((1 + z)/2, min = 0, max = 1) × 255.

  • Table 7 reports average indoor mean and median angular errors of 15.549 and 9.300 across NYUv2, DIODE-indoor, and ScanNet, the lowest reported averages among the compared methods. Lotus-2-Normal’s indoor median errors are unavailable.
  • The four-dataset mean error summarized in Table 1 is 18.928 versus 19.642 for Lotus-2. On VKitti alone, Vision Banana is worse: 29.063 mean and 10.699 median versus 28.894 and 9.677.
  • Figure 8 shows finer visible detail in selected outputs than Lotus-2-Normal, whose predictions came from https://huggingface.co/spaces/haodongli/Lotus-2_Normal. The paper notes that Lotus-2 was trained on Virtual KITTI 2 whereas Vision Banana was not.

Vision Banana has the lowest reported average indoor mean and median errors, but different methods lead individual dataset metrics, and Lotus-2-Normal’s indoor medians are unavailable. Vision Banana’s VKitti errors are slightly higher than Lotus-2-Normal’s.

Surface-normal angular errors on three indoor datasets and one outdoor dataset.
Surface-normal angular errors on three indoor datasets and one outdoor dataset.

The cat and turtle examples retain more local detail in Vision Banana’s normal maps than in the displayed Lotus-2 outputs. The outdoor example illustrates why apparent sharpness and benchmark angular accuracy should be assessed separately.

Selected surface-normal predictions from Lotus-2-Normal and Vision Banana.
Selected surface-normal predictions from Lotus-2-Normal and Vision Banana.

4. Discussion

The results support RGB image generation as a shared interface for several dense perception tasks while retaining assessed synthesis and editing capabilities. The authors also argue that generative modeling resolves ambiguous targets by selecting a coherent output mode, but the experiments do not isolate that mechanism from pretraining scale, data, or instruction tuning.

  • The claims about a paradigm shift and Artificial General Intelligence from Vision are the authors’ interpretations, not experimentally established outcomes.
  • Evaluation is restricted to monocular images. Broader task mixtures, multi-view inputs, video inputs and generators, and integration with language models are proposed future directions.
  • The authors acknowledge significantly greater computational overhead than lightweight specialists; the PDF supplies no latency or cost comparison.

Supplementary evidence separates instance-mask quality from target-presence classification and illustrates errors in instance granularity. Table 8 provides the subset breakdown, while Figures 9 and 10 compare predictions with three independent annotations and distinguish plausible-looking masks from benchmark-matched masks.

  • The average pmF1 advantage over DINO-X is small, 56.0 versus 55.2, whereas the cgF1 advantage is much larger, 47.5 versus 21.3. The authors acknowledge Gemini’s contribution to the strong IL_MCC result.
  • Figure 9 includes 17 nail masks and 3 lightbulb masks, with best F1 scores of 0.72 and 0.60. The authors attribute part of the scoring gap to smaller or finer masks differing from SAM-produced annotations; this is a qualitative explanation rather than a controlled bias analysis.
  • Figure 10 shows 28 post office boxes merged into one mask and a requested crowd split into 24 masks. Figures 11 and 12 illustrate retained generation and editing, although the second displayed editing instruction does not match its tea-table images.

The seven-subset breakdown separates presence classification from positive-query mask quality. Vision Banana’s average pmF1 of 56.0 narrowly exceeds DINO-X’s 55.2, but DINO-X leads that metric on several subsets despite its lower average gated score. Comparator results are taken from the SAM 3 paper; the authors evaluate their pipeline with the official SA-Co codebase.

Per-subset SA-Co/Gold instance-segmentation results for benchmark-trained and zero-shot methods.
Per-subset SA-Co/Gold instance-segmentation results for benchmark-trained and zero-shot methods.

The examples show that visually plausible masks can score poorly under strict IoU thresholds. Vision Banana predicts 17 nail masks with best F1 of 0.72 and three lightbulb masks with 0.60; the claimed annotation-bias explanation remains qualitative. The row labeled “eggs Benedict” scores 0.00 despite being included among the selected success cases.

Selected instance-segmentation success cases compared with three ground-truth annotations and SAM 3 predictions.
Selected instance-segmentation success cases compared with three ground-truth annotations and SAM 3 predictions.

The failure cases expose incorrect instance granularity: 28 post office boxes collapse into one mask, while a crowd is divided into 24 masks. Missing wall regions and merged furniture also show errors in target coverage and separation.

Instance-segmentation failure cases involving merging, group splitting, missed targets, and ambiguous boundaries.
Instance-segmentation failure cases involving merging, group splitting, missed targets, and ambiguous boundaries.

The beach, vehicle recoloring, and office-background examples illustrate retained editing capabilities. The second displayed instruction refers to a shelf and picture frame although its images show a tea-table scene, so that row cannot establish compliance with the printed instruction.

Original images and paired editing outputs from Vision Banana and Nano Banana Pro for ImgEdit examples.
Original images and paired editing outputs from Vision Banana and Nano Banana Pro for ImgEdit examples.

Appendix

  • A. Multi-stage clustering algorithm for parsing instance masks addresses color drift, generation noise, and boundary halos. It initializes background pixels with color tolerance τ = 14, performs seed-based floodfill with 16-connectivity, removes components smaller than θsize = 2 × 10−4 of image area, and prunes components whose preserved-area ratio after 3 × 3 erosion is below θerosion = 0.1. Similar-color components can then merge if their combined bounding-box area is no more than γ = 5.0 times the sum of their individual bounding-box areas.
  • B. Detailed SA-Co/Gold Quantitative Evaluation reports Metaclip, SA-1B, Crowded, Food&Drink, Sports Equip., Attributes, and Wiki-Common results. Positive Micro-F1 evaluates masks over 10 IoU thresholds in [0.5 : 0.05 : 0.95], IL_MCC evaluates noun-phrase presence classification, and Equation (7) defines (cgF_1=100\cdot pmF_1\cdot IL_MCC), with pmF1 expressed as a fraction rather than the table’s percentage-scale values. The combined method leads the listed zero-shot methods on average cgF1 and pmF1, but not on pmF1 in every subset; the reported average cgF1 is not the product of the displayed average component metrics.
  • C. Qualitative Instance Segmentation Analysis compares predictions with three independent ground-truth annotations and reports each prediction’s best annotation-matched F1, averaged over the benchmark’s ten IoU thresholds. The authors discuss annotation misalignment, coherent non-overlapping masks, incorrect instance merging, splitting of requested groups, and missed objects. Crowded and cluttered scenes remain substantive weaknesses, not merely scoring artifacts.
  • D. Image generation and editing capabilities presents paired Nano Banana Pro and Vision Banana outputs for GenAI-Bench and ImgEdit prompts. These qualitative comparisons complement the main human preference results and support retention of detailed synthesis and instruction-based editing, without establishing statistical equivalence between checkpoints.

Brief Thoughts

The main contribution is a demonstration that one RGB generation interface can support competitive segmentation and geometric prediction while retaining assessed synthesis capabilities. The depth results are particularly informative because task adaptation uses synthetic supervision and inference requires no camera parameters.

The evidence is narrower than the title’s general claim: the study evaluates one base generator, provides no controlled pretraining or tuning ablations, and leaves adaptation quantities and costs unspecified. External-model assistance and instance-level failures also limit conclusions about autonomous reasoning and uniformly superior generalist perception.