Gorio Tech Blog search

Monthly AI Papers: January 2026

|

Contents

Read this article in Korean

The selected research examines how language-model capability depends on the organization of computation. Recursive Language Models separates persistent input and intermediate state from bounded neural context. Ministral 3 reuses pretrained weights through progressive pruning and distillation. mHC changes how parallel residual streams exchange information across depth. These approaches intervene at different stages—inference, model construction, and pretraining architecture—and their results should be assessed within those distinct settings. 1, 2, 3

Efficiency claims require accounting beyond context length, parameter count, or FLOPs. RLM’s external storage supports million-token inputs, but sequential sub-calls produce long-tailed latency and cost. Ministral’s reported child-training horizons exclude the parent’s original pretraining cost, and its multi-sample reasoning scores are not single-generation accuracy. mHC explicitly addresses memory traffic and pipeline communication alongside numerical stability. Together, these results support evaluating capability and resource use under stated task, sampling, and implementation conditions rather than treating nominal model specifications as complete efficiency measures.

Top Papers in January 2026

1. Recursive Language Models

  • https://arxiv.org/abs/2512.24601
  • Recursive Language Models | Summary
  • Recursive Language Models (RLM) is an inference framework that stores the prompt in a persistent programming environment. The root model writes code to inspect input fragments, delegate semantic work through programmatic model calls, and assemble outputs in variables. This preserves a string-input, string-output interface without requiring the entire prompt or every intermediate result to enter the root model’s context.
  • On BrowseComp-Plus inputs spanning 6M–11M tokens, depth=1 RLM(GPT-5) achieves 91.3% correct answers at an average API cost of $0.99 ± $1.22 per task. On 32K-token OOLONG-Pairs, it reaches 58.0% F1 versus 0.1% for direct GPT-5. GPT-5 experiments use GPT-5 at the root and GPT-5-mini for sub-calls; direct BrowseComp-Plus results encounter context limits.
  • Ablations distinguish external context access from recursion. GPT-5’s OOLONG-Pairs F1 rises from 43.9% without sub-calls to 58.0%, 65.5%, and 76.0% at depths 1, 2, and 3. Deeper recursion is not uniformly beneficial: Qwen3-Coder’s OOLONG score falls from 48.0 at depth=1 to 26.0 at depth=2, while context-offloaded OpenCode outperforms the GPT-5 RLM variants on BrowseComp-Plus.


2. Ministral 3

  • https://arxiv.org/abs/2601.08584
  • Ministral 3 | Summary
  • Ministral 3 releases nine Apache 2.0 open-weight dense models: Base, Instruct, and Reasoning variants at 3B, 8B, and 14B scales, all supporting image understanding. Cascade Distillation follows a prune-distill-repeat sequence from the 24B Mistral Small 3.1 parent through 14B, 8B, and 3B. Smaller checkpoints provide initialization weights, while the original parent remains the pretraining logit teacher.
  • The 14B Base model retains substantial parent capability: MMLU-Redux is 82.0 versus 82.7, MBPP matches at 71.6, and MATH improves to 67.6 versus 55.8. Retention is task-dependent. NaturalQS decreases as models shrink, and MathVista falls from the parent’s 51.3 to 43.6, 35.7, and 23.3 for the 14B, 8B, and 3B children.
  • Ministral 3 14B Reasoning exceeds Qwen 3 14B on all six reported reasoning benchmarks, including AIME 2025 at 85.0 versus 73.7 and LiveCodeBench v6 at 64.6 versus 59.3. The report states pass@16 evaluation except pass@5 for LiveCodeBench. At 8B, comparisons are mixed: Ministral ties Qwen3-VL on AIME 2024, leads on LiveCodeBench, and trails on the other four benchmarks.


3. mHC: Manifold-Constrained Hyper-Connections

  • https://arxiv.org/abs/2512.24880
  • mHC: Manifold-Constrained Hyper-Connections | Summary
  • mHC: Manifold-Constrained Hyper-Connections addresses signal amplification in Hyper-Connections’ expanded residual streams. It uses Sinkhorn-Knopp normalization to approximate non-negative, doubly stochastic residual mixing matrices. Exact double stochasticity conserves stream averages and makes residual mixing non-expansive while permitting information exchange; it does not make every representation unchanged or establish stability of the full nonlinear network.
  • The published scaling experiments use 3B, 9B, and 27B MoE models with residual expansion rate 4 and 20 Sinkhorn-Knopp iterations. On the 27B model, mHC lowers final training loss by 0.021 relative to standard residual connections and improves all eight downstream benchmarks. It exceeds HC on seven; MATH 4-shot exact match is the exception, at 26.0 versus HC’s 26.4 and the baseline’s 22.0.
  • Measured composite backward residual propagation gain peaks near 1.6 for mHC rather than nearly 3000 for HC. This selected-sequence row- and column-sum diagnostic is not a full-network Jacobian bound, and finite normalization iterations leave approximation error. Fused mixed-precision kernels, selective recomputation, and an extended DualPipe schedule address widened-state overhead; the authors report 6.7% additional training time in in-house large-scale training at expansion rate 4, without a hardware-specific timing breakdown establishing portability.


Sources

[1] Recursive Language Models

[2] Ministral 3

[3] mHC: Manifold-Constrained Hyper-Connections