Gorio Tech Blog search

Monthly AI Papers: April 2026

|

Contents

Read this article in Korean

Generative backbones are being adapted into systems with measurable perception and spatial outputs. Vision Banana encodes segmentation, depth, and surface normals as decodable RGB images while retaining image synthesis and editing. Lyra 2.0 generates camera-controlled video and reconstructs it into explicit 3D Gaussians and meshes. These studies evaluate generation through perception benchmarks and reconstructed-scene metrics, extending assessment beyond visual plausibility. 1, 2

Consistency is addressed through explicit access to prior observations or causal predictions. Lyra 2.0 retrieves history images using per-frame geometry and trains recovery from imperfect generated histories. I-DLM jointly trains masked proposals and clean causal predictions, then overlaps proposal verification with further generation. Their results support these combined designs, without establishing either mechanism as the sole explanation for the gains. 3

Efficiency gains remain tied to operating conditions and separate quality objectives. Lyra 2.0’s distilled model accelerates segment generation but loses camera control and style consistency on Tanks and Temples. I-DLM reports higher concurrent throughput under fixed-length burst workloads, while larger trained strides reduce some benchmark scores. Vision Banana reports greater computational overhead than lightweight perception specialists without supplying latency or cost comparisons.

Top Papers in April 2026

1. Image Generators are Generalist Vision Learners

  • https://arxiv.org/abs/2604.20329
  • Image Generators are Generalist Vision Learners | Summary
  • Image Generators are Generalist Vision Learners instruction-tunes Nano Banana Pro into Vision Banana by mixing a low proportion of vision-task data into the original generation mixture. Segmentation, metric depth, and surface normals share a decodable RGB output interface, avoiding task-specific architectures and custom task losses. Exact mixing ratios, example counts, tuning duration, and compute are not reported.
  • On Cityscapes, Vision Banana reaches 69.9 mIoU versus SAM 3’s 65.2 among the listed zero-shot transfer methods; benchmark-trained SegMan-L reaches 84.2. RefCOCOg UMD val yields 73.8 cIoU. Here, zero-shot transfer excludes training on the benchmark’s in-domain split, not task supervision generally, and does not establish the contents of base-model pretraining.
  • Two segmentation results measure assisted pipelines: ReasonSeg val reaches 79.3 gIoU with Gemini 2.5 Pro translating reasoning queries into descriptive references; SA-Co/Gold reaches 47.5 cgF1 with Gemini 3.1 Flash-Lite filtering object presence. These scores should not be attributed to standalone reasoning or presence detection by Vision Banana.


2. Lyra 2.0: Explorable Generative 3D Worlds

  • https://arxiv.org/abs/2604.13036
  • Lyra 2.0: Explorable Generative 3D Worlds | Summary
  • Lyra 2.0: Explorable Generative 3D Worlds expands a single image along a user-defined camera trajectory through a retrieve–generate–update loop. Spatial memory stores independent per-frame geometry rather than a fused global reconstruction. Retrieved images supply appearance, while depth-warped canonical coordinates guide attention through geometric correspondences. FramePack and self-augmentation address temporal drift.
  • The full model leads the evaluated baselines on principal appearance and style metrics across DL3DV and Tanks and Temples, but GEN3C retains better camera control and lower reprojection error. On DL3DV, Lyra 2.0 scores 64.67 for Camera Controllability and 0.076 reprojection error, versus GEN3C’s 69.54 and 0.068. Improved appearance therefore does not imply uniformly better geometry or control.
  • A reconstruction model adapted from Depth Anything v3 produces compact 3D Gaussians and surface meshes. Downsampling its Gaussian head with k = 2 reduces Gaussian count by 4×. Fine-tuning on generated videos improves every listed reconstructed-scene metric on both datasets; on Tanks and Temples, LPIPS-P falls from 0.409 to 0.372. LPIPS-P is a reconstruction-based consistency proxy, not ground-truth geometric accuracy.


3. Introspective Diffusion Language Models

  • https://arxiv.org/abs/2604.11035
  • Introspective Diffusion Language Models | Summary
  • Introspective Diffusion Language Models converts pretrained autoregressive models using causal attention, next-token logit shifting, and supervision on masked and clean positions. Introspective Strided Decoding (ISD) verifies previous proposals and generates new ones in the same forward pass, enabling causal KV-cache reuse without a separate draft model. Conversion uses 4.5B tokens of generated reasoning responses and training on 8 H100 GPUs.
  • Strict ISD acceptance and rejection correction preserve the converted model’s causal anchor distribution. Residual ISD (R-ISD) activates gated LoRA only at masked proposal positions, leaving clean verification positions on unchanged base weights and preserving the original AR target distribution under the stated causal execution. Distribution preservation alone does not guarantee identical sampled sequences across arbitrary random-number or numerical execution choices.
  • I-DLM-8B scores 69.6 on AIME-24 and 45.7 on LiveCodeBench-v6, versus 43.3 and 30.4 for LLaDA-2.1-mini (16B). Against Qwen3-8B, results are mixed: MATH-500 rises from 95.8 to 96.8, while AIME-25 falls from 65.4 to 60.8 and LiveCodeBench-v6 from 50.3 to 45.7. Missing uncertainty estimates and mixed reproduced versus published baselines limit broader parity claims.


Sources

[1] Image Generators are Generalist Vision Learners

[2] Lyra 2.0: Explorable Generative 3D Worlds

[3] Introspective Diffusion Language Models