Gorio Tech Blog search

Attention Residuals | Summary

|

Contents

This article explains the key points of Attention Residuals.

  • 2026-03-16 (arXiv)
  • Kimi Team, Chen, Guangyu, Zhang, Yu, Su, Jianlin, Xu, Weixin, Pan, Siyuan, Wang, Yaoyu, Wang, Yucheng, Chen, Guanduo, Yin, Bohong, et al.
  • Paper
  • Github

Read this article in Korean


Summary

  • Attention Residuals (AttnRes) replaces fixed residual addition with softmax-weighted retrieval over earlier layer outputs. A learned pseudo-query per layer produces token-dependent weights through input-dependent keys, allowing each layer to select its own depth-wise mixture.
  • Block AttnRes compresses groups of layer outputs into summed block representations, reducing retained-history storage and cross-stage representation transfer from O(Ld) to O(Nd). With cross-stage caching and two-phase inference, the report measures less than 4% training overhead under pipeline parallelism and less than 2% inference latency overhead on typical workloads.
  • Five-size scaling experiments show consistently lower validation loss, with a fitted 1.25× compute advantage for Block AttnRes. A Kimi Linear model with 48B total / 3B activated parameters trained on 1.4T tokens improves 14 of 15 reported benchmarks and ties the remaining one, including GPQA-Diamond 36.9 → 44.4 and HumanEval 59.1 → 62.2. The evaluation is confined to the studied architecture family and provides no repeated-run uncertainty estimates.

1 Introduction

Standard residual connections provide gradient highways and define a fixed depth-wise aggregation rule: unrolling the recurrence reveals an equally weighted sum of the embedding and all preceding outputs. The introduction links this aggregation to PreNorm dilution and proposes content-aware retrieval across depth rather than another modification of the additive recurrence.

  • The contributions comprise Full and Block AttnRes, a structured-matrix interpretation of residual mechanisms, infrastructure optimizations, and scaling, ablation, and downstream evaluations.
  • The source repository is https://github.com/MoonshotAI/Attention-Residuals.

The diagrams isolate the aggregation change: standard residuals merge outputs into one stream, Full AttnRes retains independently selectable earlier outputs, and Block AttnRes selects summed groups. The storage reduction comes from coarser source granularity, not from removing learned depth-wise weighting.

Standard residual connections, Full Attention Residuals, and Block Attention Residuals, with retained-history memory O(Ld) versus O(Nd).
Standard residual connections, Full Attention Residuals, and Block Attention Residuals, with retained-history memory O(Ld) versus O(Nd).

2 Motivation

Notation uses input batches of shape B × T × d, with batch size B, sequence length T, and hidden dimension d. For a single token, h_l is the state entering layer l, h_1 is the embedding, and f_l is the layer transformation. Each attention or MLP sublayer counts as an individual layer, so a Transformer block contains two such layers.

2.1 Training Deep Networks via Residuals

Residual Learning uses h_l = h_{l−1} + f_{l−1}(h_{l−1}), which expands into the embedding plus all previous layer outputs. Expanding the backward product of factors I + ∂f_j/∂h_j preserves an identity term, providing a direct gradient path. Highway networks generalize the update with element-wise gates but still carry a single aggregated state.

  • Generalizing Residuals writes h_l = α_l h_{l−1} + β_l f_{l−1}(h_{l−1}); standard residuals set both coefficients to 1, whereas Highway uses α_l = 1 − g_l and β_l = g_l, with element-wise multiplication for vector gates.
  • Limitations identifies three consequences of single-state aggregation: no selective access to individual earlier outputs, inability to recover information lost during aggregation, and increasingly large later-layer outputs needed to influence the accumulated state.

3 Attention Residuals: A Unified View of Time and Depth

The Duality of Time and Depth compares residual accumulation across layers with RNN recurrence across tokens. AttnRes replaces this compressed-state interface with (h_l = \alpha_{0\to l}h_1 + \sum_{i=1}^{l-1}\alpha_{i\to l}f_i(h_i)), where the source weights sum to 1. The authors argue that typical depth L < 1000 makes quadratic depth attention computationally feasible.

3.1 Full Attention Residuals

Full AttnRes sets q_l = w_l and k_i = v_i, with v_0 = h_1 and v_i = f_i(h_i) for earlier layers. It uses α_{i→l} = exp(w_l^T RMSNorm(k_i)) / Σ_{j=0}^{l−1} exp(w_l^T RMSNorm(k_j)) and forms the next input from unnormalized values. Key normalization prevents source magnitude from directly dominating selection.

  • The pseudo-query is a learned d-dimensional parameter; the weights are nevertheless input-dependent because the keys are token-specific layer outputs.
  • Per token, arithmetic is O(L²d) and retained history is O(Ld). These outputs already exist for backpropagation in vanilla training, but activation recomputation and pipeline parallelism introduce additional preservation and transfer costs.
  • Blockwise optimization batches fixed queries against already available history before sequential layer execution, then handles within-block dependencies separately. It reduces amortized per-layer I/O to O((S + N)d), but does not eliminate Full AttnRes’s O(Ld) cross-stage representation transfer.

3.2 Block Attention Residuals

Block AttnRes partitions L layers into N blocks and represents completed block n by b_n = Σ_{j∈B_n} f_j(h_j). Each layer attends over the embedding b_0, completed block sums, and—after the first layer of its block—the evolving intra-block partial sum. The final output aggregates the completed blocks with the embedding retained as a source.

  • Intra-Block Accumulation sums transformation outputs, not their attention-weighted inputs. The uniform-block derivation assumes S = L/N; for an incomplete final block, its final partial sum becomes its representation. RMSNorm normalizes keys from both complete blocks and partial sums.
  • Inter-Block Attention preserves distant history but cannot select individual outputs inside a completed block. Figure 2 shows separate AttnRes operations before attention and MLP.
  • Retained representations require O(Nd) storage. Section 3.2 states O(N²) computation, whereas Related Work gives O(LN) for aggregation across L layers; the latter accounts for each layer attending over block sources. N = L recovers Full AttnRes, while N = 1 retains a separately weighted embedding and partial sum rather than reproducing the original unnormalized residual addition.
  • Approximately eight blocks recover most of the gains in the scaling experiments. A fixed block count also bounds the depth-history cache used during inference.

The pseudocode normalizes keys, computes a softmax across depth sources, and mixes unnormalized values. Separate calls before attention and MLP allow different sublayer mixtures, while partial_block accumulates transformation outputs and is added to completed history at a block boundary.

PyTorch-style pseudocode maintaining completed block history and the current intra-block partial sum.
PyTorch-style pseudocode maintaining completed block history and the current intra-block partial sum.

4 Infrastructure Design

Infrastructure Design addresses redundant transfers between pipeline stages, repeated history reads during inference, and representation storage during long-context prefilling. The proposed solutions are cross-stage caching, batched inter-block attention followed by sequential intra-block merging, and sequence-sharded prefilling.

4.1 Training

Training caches previously received block representations on each physical pipeline rank so later virtual stages transfer only incremental history. With P physical stages, V virtual stages, C = PV chunks, and average block production N_p under the report’s simplifying assumption, Comm_naïve = C(C−1)N_p d/2 and Comm_cached = P(P−1)N_p d/2 + (V−1)P²N_p d.

  • Without caching, each transition resends the growing history. Cross-stage caching reduces peak per-transition scaling from O(C) to O(P), a V× improvement, and also applies to backward communication.
  • The P = 4, V = 2 example eliminates 6 redundant block transmissions in the second virtual stage. Block boundaries need not align with physical stage boundaries.
  • Each block is stored once across virtual stages, and activation checkpointing removes inter-block attention intermediates while retaining an input with the same size as the replaced hidden state. The report states negligible training overhead without pipeline parallelism and less than 4% measured end-to-end overhead with it.

The second virtual stage reuses history already cached on each rank and transfers only incremental blocks. With 4 physical ranks and 2 virtual stages, the example removes 6 redundant block transmissions; hatched boundaries also show that AttnRes blocks can span physical stages.

Cache-based pipeline communication with 4 physical ranks, 2 virtual stages per rank, and incremental block transfers.
Cache-based pipeline communication with 4 physical ranks, 2 virtual stages per rank, and incremental block transfers.

4.2 Inference

Inference exploits pseudo-queries that are independent of hidden states. Phase 1 batches all S queries in a block against completed history; Phase 2 handles each layer’s evolving intra-block source sequentially and combines the two results through exact online-softmax merging.

  • Algorithm 1 retains attention numerators and normalization statistics, rescales both phases to a shared maximum, and divides the combined numerator by the combined normalizer. This is equivalent to a joint softmax over both source sets, rather than an approximation.
  • For Block AttnRes, amortized per-layer reads are (N/S + 3)d and writes are 2d. Under L = 128, N = 8, S = 16, total residual I/O is 5.5d, compared with 3d for standard residuals, 24d for Full AttnRes, and 34d for mHC with m = 4; these estimates exclude the internal layer transformation.
  • Memory-efficient prefilling shards N × T × d representations over tensor-parallel devices and merges locally within a reduce-scatter/all-gather path. The report’s 128K-token, 8-block example falls from 15 GB to roughly 1.9 GB per device, and below 0.3 GB with 16K chunked prefill.
  • The reported inference latency overhead is less than 2% on typical workloads. The report does not supply a detailed hardware and workload latency sweep.

Under L = 128, N = 8, S = 16, and m = 4, Block AttnRes has 5.5d residual I/O per token per layer, versus 3d for standard residuals, 24d for Full AttnRes, and 34d for mHC. These amortized traffic estimates exclude the layer function and do not imply corresponding end-to-end latency ratios.

Per-token, per-layer residual-mechanism reads, writes, and total memory I/O for standard residuals, mHC, and two-phase AttnRes.
Per-token, per-layer residual-mechanism reads, writes, and total memory I/O for standard residuals, mHC, and two-phase AttnRes.

5 Experiments

Architecture Details keeps the Kimi Linear MoE architecture unchanged except for residual aggregation: Kimi Delta Attention (KDA) and Multi-Head Latent Attention (MLA) appear in a 3:1 ratio, each followed by an MoE feed-forward layer. AttnRes adds one RMSNorm and one pseudo-query vector per sublayer, contributing a negligible fraction of total parameters.

  • All pseudo-queries are initialized to zero, producing uniform source weights and an equal-weight average at initialization. The authors identify this initialization as necessary to avoid training volatility in their experiments.

5.1 Scaling Laws

Scaling Laws compares PreNorm, Full AttnRes, and Block AttnRes with approximately eight blocks across five MoE sizes: 194M, 241M, 296M, 436M, and 528M activated parameters excluding embeddings. All use an 8192-token context, a cosine learning-rate schedule, and identical within-size hyperparameters selected for the baseline.

  • Training token counts are respectively 38.7B, 45.4B, 62.1B, 87.9B, and 119.0B. Corresponding Baseline / Block / Full validation losses are 1.931 / 1.909 / 1.899; 1.895 / 1.875 / 1.874; 1.829 / 1.809 / 1.804; 1.766 / 1.746 / 1.737; and 1.719 / 1.693 / 1.692.
  • The fitted curves are ℒ = 1.891 × C^(−0.057) for Baseline, ℒ = 1.870 × C^(−0.058) for Block AttnRes, and ℒ = 1.865 × C^(−0.057) for Full AttnRes, with C measured in PFLOP/s-days.
  • At 5.6 PFLOP/s-days, the report gives fitted Block loss 1.692 versus Baseline 1.714 and states a 1.25× compute advantage. This fitted comparison is distinct from the largest measured row in Table 2.
  • The Full–Block gap is 0.001 at the largest measured size. Table 2’s mHC(-lite) reference has the lowest loss in the 241M row, so Full AttnRes does not outperform it at every size.

Both AttnRes variants lower validation loss relative to PreNorm at all five measured sizes, and the Full–Block gap is 0.001 at 528M activated parameters. The mHC(-lite) reference has the lowest loss in the 241M row, qualifying the surrounding prose's broad superiority claim.

Model configurations, training hyperparameters, token counts, and validation losses across five activated-parameter scales.
Model configurations, training hyperparameters, token counts, and validation losses across five activated-parameter scales.

5.2 Main Results

Training recipe uses 27 Transformer blocks, or 54 attention/MLP layers, with 8 of 256 routed experts and 1 shared expert, yielding 48B total / 3B activated parameters. Block AttnRes groups 6 layers per block, giving 9 completed blocks plus the embedding, or 10 final depth-wise sources.

  • The matched recipe uses a 4096-token context, Muon optimizer, WSD (Warmup–Stable–Decay) learning-rate schedule, and global batch size of 8M tokens.
  • Training comprises 1T pre-training tokens followed by approximately 400B high-quality mid-training tokens under the Moonlight annealing recipe. Subsequent context-extension training reaches 32K tokens.
  • The hybrid KDA/MLA architecture uses MLA without positional encodings (NoPE), so the reported context extension requires neither YaRN nor attention-temperature rescaling.

Training dynamics over the 1T-token phase show lower validation loss for Block AttnRes throughout training, with a wider gap during decay. Figure 5 also shows a bounded periodic output-magnitude pattern and more evenly distributed gradients across depth, whereas baseline outputs increase sharply in later blocks and baseline gradients are largest in early blocks.

  • The authors interpret the periodic magnitude pattern as accumulation confined within blocks and selective aggregation at block boundaries.
  • These measurements support mitigation of PreNorm dilution in the studied model, but do not establish a general stability guarantee.

The loss panel favors Block AttnRes throughout the 1T-token training phase. Its output magnitudes remain low with a periodic pattern, and its gradient profile is flatter than the baseline's; these observations support mitigation of dilution in this model rather than a general stability guarantee.

Training validation loss, final output magnitudes, and gradient magnitudes across Transformer blocks.
Training validation loss, final output magnitudes, and gradient magnitudes across Transformer blocks.

Downstream performance follows the Kimi Linear evaluation protocol and compares models trained with the same recipe. Block AttnRes improves 14 reported benchmarks and ties MMLU-Pro; GPQA-Diamond has the largest listed gain. The proposed connection between depth-wise retrieval and compositional reasoning remains a hypothesis.

  • General results, Baseline → AttnRes: MMLU 73.5 → 74.6; MMLU-Pro 52.2 → 52.2; GPQA-Diamond 36.9 → 44.4; BBH 76.3 → 78.0; ARC-Challenge 64.6 → 65.7; HellaSwag 83.2 → 83.4; TriviaQA 69.9 → 71.8.
  • Math & Code results: GSM8K 81.7 → 82.4; MGSM 64.9 → 66.1; Math 53.5 → 57.1; CMath 84.7 → 85.1; HumanEval 59.1 → 62.2; MBPP 72.0 → 73.9.
  • Chinese results: CMMLU 82.0 → 82.9 and C-Eval 79.6 → 82.5. No confidence intervals or repeated-run variability are reported for these differences.

5.3 Ablation Study

Comparison with prior methods uses the 16-head configuration from Table 2 under identical hyperparameters and compute budgets. Validation loss is 1.766 for PreNorm, 1.767 for DenseFormer, 1.747 for mHC, 1.737 for Full AttnRes, and 1.746 for Block AttnRes with S = 4. These comparisons favor content-dependent cross-layer aggregation in this setting, although Block AttnRes’s margin over mHC is only 0.001.

  • Cross-layer access extends beyond a short window: sliding-window aggregation over the embedding plus 8 recent outputs reaches only 1.764.
  • Figure 6 reports S = 32, 16, 8, 4, 2 losses of 1.757, 1.753, 1.748, 1.746, 1.746; Full AttnRes, S = 1, reaches 1.737. Coarse blocks retain some gains, but fine-grained access remains better.
  • Table 4 and Figure 6 call this a 16-layer model, while Table 2 and Figure 8 identify 16 Transformer blocks containing 16 attention and 16 MLP layers. The captions therefore use a different layer-count convention from the paper’s sublayer-based notation.

Full AttnRes reaches 1.737 versus PreNorm's 1.766, while static mixing, sigmoid weighting, and removing RMSNorm each worsen its loss. An input-dependent query reaches 1.731 but adds projection and sequential-access costs; Block multihead aggregation also underperforms the single-mixture Block design.

Validation-loss ablations of residual alternatives, cross-layer access, queries, normalization, and depth-attention design.
Validation-loss ablations of residual alternatives, cross-layer access, queries, normalization, and depth-attention design.

Reducing block size from 32 to 4 lowers loss from 1.757 to 1.746, with no further measured gain at size 2. Full AttnRes at size 1 remains better at 1.737, exposing a remaining cost of block compression despite its favorable storage and communication trade-off.

Validation loss across Block AttnRes block sizes, with standard residual and Full AttnRes reference lines.
Validation loss across Block AttnRes block sizes, with standard residual and Full AttnRes reference lines.

Component design ablations favor softmax competition, normalized keys, and content-dependent weights. An input-dependent query lowers Full AttnRes loss further to 1.731, but the authors retain the fixed learned query because a d × d projection and sequential query generation add parameter and decoding-access costs.

  • Input-independent learned mixing raises Full loss to 1.749; sigmoid weighting gives 1.741; removing key RMSNorm gives 1.743.
  • For Block AttnRes, H = 16 multihead depth aggregation gives 1.752 versus 1.746, and removing RMSNorm gives 1.750.
  • The multihead result is consistent with whole-representation mixtures being sufficient in this experiment, but does not establish that channel-wise specialization is generally unhelpful.

5.4 Analysis

5.4.1 Optimal Architecture studies 25 configurations under approximately 6.5 × 10^19 FLOPs and 2.3 × 10^8 activated parameters, with d_ff/d_model ≈ 0.45. The grid spans d_model/L_b ∈ {15, 30, 45, 60, 75} and H/L_b ∈ {0.3, 0.4, 0.5, 0.6, 0.7}, where L_b = L/2.

  • AttnRes lowers loss in every grid cell by 0.019–0.063. Both optima have H/L_b ≈ 0.3, but the baseline optimum is d_model/L_b ≈ 60 with loss 1.847, versus ≈ 45 and 1.802 for AttnRes.
  • At fixed parameter budget, the lower width-to-depth ratio indicates a deeper, narrower preferred configuration. The authors caution that greater sequential depth can increase inference latency, so this diagnostic is not a deployment recommendation.

AttnRes lowers loss in every fixed-budget architecture cell. Its optimum shifts from the baseline's width-to-depth ratio 60 to 45 at the same head-to-depth ratio 0.3, indicating a deeper, narrower preferred allocation; the loss sweep does not measure the latency cost of additional sequential depth.

A 5 × 5 architecture sweep under approximately 6.5 × 10^19 FLOPs and 2.3 × 10^8 activated parameters.
A 5 × 5 architecture sweep under approximately 6.5 × 10^19 FLOPs and 2.3 × 10^8 activated parameters.

5.4.2 Analyzing Learned AttnRes Patterns examines token-averaged weights for a model with 16 attention and 16 MLP layers, comparing Full with N = 8 Block AttnRes. Both retain strong local pathways while assigning persistent weight to the embedding and selective weight to distant sources.

  • Preserved locality appears as diagonal dominance, with off-diagonal concentrations indicating learned skip connections.
  • Layer specialization appears in broader pre-attention mixtures and sharper pre-MLP reliance on recent sources; embedding weight is especially visible before attention.
  • Block AttnRes preserves these qualitative patterns with sharper distributions. Because the heatmaps average over tokens, they do not show the extent of token-by-token routing variation or establish the authors’ proposed implicit-regularization effect.

Diagonal concentrations show that local sources remain important, while embedding weights and off-diagonal peaks show persistent and selective long-range depth access. Pre-attention mixtures are broader than pre-MLP mixtures, and Block AttnRes retains these patterns; token averaging hides individual routing decisions.

Token-averaged depth-attention weights before attention and MLP for Full AttnRes and N = 8 Block AttnRes.
Token-averaged depth-attention weights before attention and MLP for Full AttnRes and N = 8 Block AttnRes.

6 Discussions

Discussions organizes residual mechanisms through sequence–depth duality and structured depth-mixing matrices. Table 5 distinguishes fixed, learned-static, and input-dependent weights, and separates single-state recurrence, multi-state recurrence, and access to retained earlier outputs.

6.1 Sequence-Depth Duality

Sequence-Depth Duality connects recurrent sequence updates to recurrent depth updates. Test-Time Training writes W_t = W_{t−1} − η∇ℓ(W_{t−1}; x_t), and the linear case yields additive linear attention S_t = S_{t−1} + k_t v_t^T. Standard residuals share this additive form across layers.

  • The report associates sequence gates with Highway networks, delta-rule models with DDL, and MRLA with gated linear attention. AttnRes instead exposes earlier outputs through direct depth-wise softmax attention.
  • This is a structural analogy, not a claim that AttnRes performs Test-Time Training or removes sequential dependencies between layer transformations.

The table separates dynamic weighting from direct cross-layer access: Highway and HC/mHC have input-dependent updates but operate on carried states, whereas AttnRes selects retained earlier sources. DenseFormer's learned-static cross-layer weights show that source access and content dependence are distinct design choices.

Residual update rules grouped into single-state recurrence, multi-state recurrence, and cross-layer access.
Residual update rules grouped into single-state recurrence, multi-state recurrence, and cross-layer access.

6.2 Residual Connections as Structured Matrices

Residual Connections as Structured Matrices writes h_l = Σ_{i=0}^{l−1} M_{i→l}v_i and compares the structure of the depth-mixing matrix M. Standard residuals yield an all-ones lower-triangular matrix; scalar-gated Highway is 1-semiseparable, while m-stream HC/mHC is m-semiseparable. These semiseparable ranks describe the recurrence structure, not the ordinary rank of the full triangular matrix.

  • Highway source weights factor into a write gate and a product of later carry gates. Their sum-to-one property connects the recurrence to softmax-free stick-breaking attention.
  • HC/mHC unrolls to M_{i→l} = β_i^T A^×_{i+1→l} α_l, with cumulative transition products between source and receiving layer. mHC constrains transition matrices to be doubly stochastic.
  • Full AttnRes computes normalized softmax weights directly. Block AttnRes ties the weights of outputs inside each completed block and adds the current partial sum as a source; the report characterizes its effective rank as lying between N and N + S, without a separate derivation of that bound.

Highway and HC/mHC weights contain products of intervening recurrence operators, whereas Full AttnRes scores earlier sources directly. Block AttnRes shares scores across completed source groups, making the loss of within-block selection explicit; the displayed AttnRes scores are unnormalized.

Depth-mixing matrices for Highway, HC/mHC, Full AttnRes, and Block AttnRes with L = 4 and block size S = 2.
Depth-mixing matrices for Highway, HC/mHC, Full AttnRes, and Block AttnRes with L = 4 and block size S = 2.

Practicality uses the matrix view to identify persistent high-weight sources and connect decomposable attention kernels with recurrent implementations. Prior Residuals as Depth-Wise Linear Attention interprets HC/mHC’s α_l as a query, β_i as a key, and cumulative transitions as depth-relative operators. Multiple streams expand the recurrent state from d to d × m.

  • The report contrasts these matrix-state linear-attention mechanisms with AttnRes’s depth-wise softmax attention. This interpretation relates the designs without implying equivalent expressivity or deployment cost.
  • A factorized kernel ϕ(q, k) = φ(q)^Tφ(k) permits recurrent aggregation, whereas the chosen exponential kernel implements normalized content-selective retrieval.

Related Work distinguishes normalization and residual scaling, multi-state recurrence, and cross-layer connectivity. AttnRes combines softmax-normalized input-dependent selection over earlier outputs, a single d-dimensional pseudo-query per layer, and block compression supported by distributed-training and inference optimizations.

  • Normalization, Scaling, and Depth Stability contrasts PreNorm’s identity gradient path and magnitude growth with PostNorm’s repeated normalization along the residual path, and discusses scaled, hybrid, and gated alternatives.
  • Multi-State Recurrence includes HC/mHC, DDL, and SiameseNorm; these expand or stabilize the carried state rather than retain independently selectable earlier outputs.
  • Cross-Layer Connectivity includes static aggregation in DenseNet, ELMo, DenseFormer, and ANCRe, dynamic aggregation in MUDDFormer and MRLA, and targeted access in Value Residual Learning, LAuReL, and Dreamer.

Conclusion

Conclusion identifies learned depth-wise attention as the central change and Block AttnRes as the practical large-scale variant. Approximately eight blocks recover most of the gains within the tested scaling configurations, while Full AttnRes retains finer-grained selection at the cost of preserving and communicating earlier layer outputs.

  • Practicality at scale depends on cross-stage caching and two-phase computation, not solely on the small added parameter count. Finer blocking remains an architectural trade-off rather than an experimentally established improvement at the 48B scale.

Appendix

  • A Contributions lists contributors by significance of contribution, with project leadership appearing last. Guangyu Chen, Yu Zhang, and Jianlin Su are marked as equal contributors. This appendix provides attribution rather than additional experimental evidence.
  • B Optimized Inference I/O for Full Attention Residuals partitions execution into N blocks of S = L/N layers solely for scheduling, without compressing per-layer sources. Phase 1: Batched Inter-block Attention gives total reads dL(N−1) and writes Ld; Phase 2: Sequential Intra-block Attention adds NS(S−1)d reads and Ld writes. The resulting amortized costs are (S + N − 2)d reads, 2d writes, and (\text{Total I/O per layer}=(S+N)d).

Brief Thoughts

The controlled loss comparisons, component ablations, and large-model downstream results provide complementary evidence for AttnRes. Static mixing and short-window access perform worse than Full AttnRes, while Block AttnRes preserves much of its benefit. These comparisons support content-dependent access to distant depth history more directly than the sequence–depth analogy alone.

Generality remains unresolved: the experiments use the Kimi Linear architecture family, without repeated-run uncertainty or a detailed latency benchmark matrix. The 1.25× advantage is a fitted training-compute comparison, not a measured wall-clock speedup. Token-averaged routing plots support qualitative interpretation but do not establish the mechanism behind reasoning gains.