Gorio Tech Blog search

Monthly AI Papers: March 2026

|

Contents

Read this article in Korean

Learned allocation is a common theme across March’s architecture studies. Beyond Language Modeling uses fine-grained Mixture-of-Experts routing to allocate capacity between language and vision, improving both modalities at approximately fixed active computation. Attention Residuals instead learns token-dependent mixtures of earlier layer outputs, replacing fixed residual accumulation with selective access across depth. These results support revisiting how models distribute capacity and reuse representations, alongside increasing model size. 1, 2

Training efficiency increasingly depends on which examples receive computation. Diverse multimodal supplementation outperforms additional task-specific data in several controlled VQA and navigation comparisons. PivotRL selects intermediate demonstration states with mixed rollout outcomes and trains short, locally verified continuations rather than regenerating complete trajectories. The efficiency measures remain distinct: matched active FLOPs do not establish equal hardware cost, Attention Residuals’ 1.25× compute advantage is fitted rather than a wall-clock speedup, and PivotRL’s rollout savings exclude complete demonstration-generation and offline-profiling costs. 3

Capability retention remains a separate criterion from target-task improvement. Every tested multimodal data mixture worsens Notes perplexity relative to text-only pretraining, despite useful visual transfer. PivotRL retains average performance near its base model across eight out-of-distribution benchmarks, where same-data supervised fine-tuning produces substantial regressions, but it scores below SFT on SWE-Bench Verified. Neither result establishes universal preservation of capabilities.

Top Papers in March 2026

1. Beyond Language Modeling: An Exploration of Multimodal Pretraining

  • https://arxiv.org/abs/2603.03276
  • Beyond Language Modeling: An Exploration of Multimodal Pretraining | Summary
  • Beyond Language Modeling: An Exploration of Multimodal Pretraining separates representation, data, architecture, world-modeling, and scaling effects. It trains the Transformer backbone from scratch with Transfusion-style next-token prediction for language and flow matching for vision, while retaining frozen pretrained visual encoders and pretrained decoders. A single SigLIP 2 Representation Autoencoder gives the strongest combined understanding and generation results among the tested representations.
  • General-data supplementation can beat specialized scaling. Training with 20B VQA tokens plus 80B text tokens reaches 37.9% average VQA accuracy, versus 35.7% with 100B VQA-only tokens. For navigation, 50B NWM tokens plus 50B raw-video tokens yields absolute trajectory error 1.410 and relative pose error 0.414, versus 1.669 and 0.450 with 100B NWM-only tokens. The approximately 1% navigation-data result uses a fixed 200B-token budget and still includes domain alignment data.
  • Under matched hyperparameters, data, token counts, and FLOPs, stacked design changes improve the Transfusion baseline from PPL 15.93 and DPG 0.45 to PPL 12.49 and DPG 0.63 with modality-specific FFNs, SigLIP 2, and MoE. Switching the final configuration to x-pred raises DPG to 0.65 while worsening PPL to 12.69. IsoFLOP fits narrow the vision–language training-data exponent gap from 0.10 for dense models to 0.05 for MoE; they do not eliminate it.


2. Attention Residuals

  • https://arxiv.org/abs/2603.15031
  • Attention Residuals | Summary
  • Attention Residuals replaces fixed residual addition with softmax-weighted retrieval over earlier layer outputs. Each layer uses a learned pseudo-query and RMSNorm-normalized, token-specific keys, so its depth-wise mixture remains input-dependent. Block AttnRes stores summed groups of outputs instead of every individual output, reducing retained-history storage and cross-stage transfer in proportion to the block count rather than layer count, at the cost of losing selection within completed blocks.
  • Both Full and Block AttnRes lower validation loss relative to PreNorm across five tested MoE sizes, from 194M to 528M activated parameters excluding embeddings. At the largest measured size, losses are 1.719 for PreNorm, 1.693 for Block, and 1.692 for Full. Approximately eight blocks recover most gains in these scaling experiments; the reported 1.25× compute advantage comes from fitted loss curves.
  • A matched Kimi Linear comparison with 48B total and 3B activated parameters, trained on 1.4T tokens, improves 14 of 15 reported benchmarks and ties MMLU-Pro. GPQA-Diamond rises from 36.9 to 44.4 and HumanEval from 59.1 to 62.2. These results support the aggregation change within the studied architecture family; repeated-run uncertainty is not reported, and the proposed explanation linking depth retrieval to reasoning gains remains a hypothesis.


3. PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost

  • https://arxiv.org/abs/2603.21383
  • PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost | Summary
  • PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost converts expert trajectories into short, on-policy RL episodes starting at intermediate assistant turns. Offline profiling retains states with positive empirical reward variance and reward mean below the difficulty threshold λdiff. Domain-specific verifiers reward locally acceptable actions rather than exact reproduction of a demonstrated completion. Selection is fixed before RL, and the numerical difficulty threshold is not supplied.
  • Four separately trained domain models start from Qwen3-30B-A3B-Thinking-2507. Average in-domain improvement is +14.11 percentage points over Base, versus +9.94 for same-data SFT. Across four training runs and eight OOD benchmarks, average changes are +0.21 and −9.83 points, respectively. PivotRL beats SFT on three of four in-domain benchmarks, but SWE-Bench Verified is lower: 32.67 versus 37.40.
  • At matched SWE-Bench Verified accuracy of 32.67%, PivotRL uses approximately 133K rollout turns versus 542K for end-to-end RL, with approximately 5.5× less cumulative rollout time on the same number of compute nodes. End-to-end RL subsequently reaches higher accuracy. The comparison establishes savings at a shared accuracy target, not higher final performance or lower total training cost.


Sources

[1] Beyond Language Modeling: An Exploration of Multimodal Pretraining

[2] Attention Residuals

[3] PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost