Gorio Tech Blog search

Language Models Can Control Their Own Attention | Summary

|

Contents

This article explains the key points of Language Models Can Control Their Own Attention.

  • 2026-09-02 (arXiv)
  • Ho, Namgyu, Ahmad, Huzama, Koh, Woosung, Yun, Se-Young, Schuster, Tal, Santos, Cicero Nogueira dos.
  • KAIST AI, Google DeepMind
  • Paper

Read this article in Korean


Summary

  • Declarative Attention (DA) is a zero-shot inference protocol in which a language model emits parseable attention declarations during generation. Across 15 long-context tasks, DA reduces average attended tokens by 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B, with accuracy drops of 1.27pp and 2.75pp; its wall-clock results are roofline estimates rather than latency measurements.

1. Introduction

Long-context decoding normally reads the full KV cache at every autoregressive step. DA instead derives a mask from generated text, avoiding an auxiliary per-step selection scan except when the model deliberately enters a full-context global phase.

2. Declarative Attention

DA organizes reasoning into <global>, <focus>, and <local> spans. Global mode retains all context for navigation, focus mode retains named context segments for extraction, and local mode removes context segments while preserving the scaffold and generated response.

The figure illustrates how generated <global>, <focus>, and <local> spans produce progressively narrower segment-level masks while retaining the scaffold and generated response. In the example, attended prompt tokens fall from 25,466 in global mode to 3,435 in focus mode and 1,124 in local mode.

Prompt structure, three declarative reasoning modes, and the corresponding segment-level KV attention masks.
Prompt structure, three declarative reasoning modes, and the corresponding segment-level KV attention masks.

2.1. Prompt overview

The prompt separates an always-visible scaffold from selectively visible context. The scaffold contains the system instruction, question, and DA instruction, allowing the task and protocol to remain available under every masking mode.

2.2. Context delivery

The input is partitioned into addressable “magic chunks,” targeting 2048-token segments and preferring paragraph, newline, sentence, clause, and word boundaries. Chunks are presented in a simulated tool-use transcript, although no tools are executed at inference time.

2.3. Decode-time interventions

The state machine begins in global mode, switches on <focus magic_chunks=”K”> or <local>, and reverts to global after the matching closing tag. A vLLM hook rewrites the request’s block table at KV-cache block granularity; global attention layers are masked, while sliding-window attention and Gated DeltaNet layers are left unchanged.

3. When does DA pay off?

DA trades additional decode steps for fewer attended KV positions per step. The favorable regime is long-context, large-batch serving in which global-attention reads are memory-bound and dominate decode cost; the paper uses a roofline decomposition to characterize this trade-off.

4. Experimental settings

The experiments evaluate six Gemma and Qwen models, with Gemma-4-31B and Qwen-3.6-27B as the headline models. The 15-source suite draws from RULER, LongBench v1/v2, LooGLE, and ZeroSCROLLS; contexts above 244K tokens, or 116K for Gemma-4-E4B, are excluded and outputs are capped at 8K tokens.

The table defines the 15 evaluation sources, their domains, task types, QA provenance, and context-length distributions. It shows that the suite covers both single-span retrieval/reasoning and multi-span reasoning, including code repositories with a 1071.1K-token mean source length before filtering.

Long-context benchmark sources grouped into single-span retrieval/reasoning and multi-span reasoning tasks.
Long-context benchmark sources grouped into single-span retrieval/reasoning and multi-span reasoning tasks.

5. Results

The study compares Vanilla full causal attention, DA-no-mask (DAnm), and DA with custom masking. Accuracy is evaluated by an LLM judge using generated rubrics, while attended tokens sum the KV positions read over decoding and roofline wall-time estimates deployment-scale decode cost.

5.1. Main results

On Gemma-4-31B, DA changes average accuracy from 87.01% to 85.74% and attended tokens from 13.43M to 6.45M per sample. On Qwen-3.6-27B, accuracy changes from 85.31% to 82.56% and attended tokens from 22.54M to 15.52M; DAnm is near accuracy-neutral but increases attended tokens, isolating the effect of masking.

The table reports per-task and macro-average accuracy and attended-token totals for Vanilla, DAnm, and DA on the two headline models. It shows heterogeneous task-level outcomes despite aggregate attention reductions and documents cases where longer generations increase total attended tokens.

Accuracy and attended-token results for Vanilla, DA-no-mask, and Declarative Attention across 15 sources.
Accuracy and attended-token results for Vanilla, DA-no-mask, and Declarative Attention across 15 sources.

The normalized comparison shows that DA produces 35% more decode steps on Gemma-4-31B and 31% more on Qwen-3.6-27B, yet uses only 48% and 69% as many attended tokens. DAnm increases attended tokens to 166% and 129% of vanilla, isolating the benefit of the mask.

Relative accuracy, decode steps, and attended tokens for Vanilla, DA, and DA-no-mask on the headline models.
Relative accuracy, decode steps, and attended tokens for Vanilla, DA, and DA-no-mask on the headline models.

5.2. Model capability scaling

Relative DA accuracy improves monotonically with scale in both model families, from 29% to 99% of vanilla accuracy across Gemma-4-E4B to Gemma-4-31B and from 64% to 97% across Qwen-3.5-4B to Qwen-3.6-27B. Gemma-4-12B exceeds vanilla attended tokens because some responses exhaust the 8K generation budget, not because masking increases its per-step attention.

The figure shows improving relative accuracy under DA as model size increases in both Gemma and Qwen families. It also attributes Gemma-4-12B’s 183% attended-token result to non-termination: excluding responses that hit the generation limit lowers the value to 98%.

Relative accuracy and attended tokens under DA across Gemma 4 and Qwen 3.5/3.6 model sizes.
Relative accuracy and attended tokens under DA across Gemma 4 and Qwen 3.5/3.6 model sizes.

5.3. Context-length scaling

For Gemma-4-31B, DA stays within roughly one percentage point of vanilla through 32K-token contexts and falls to about 96% of vanilla in the 64–256K bin. Its absolute attended-token saving increases from about 1M at short contexts to about 21M in the longest bin.

For Gemma-4-31B, DA remains near vanilla accuracy through 32K tokens before declining at longer contexts, whereas DAnm does not show the same decline. The absolute attended-token saving grows with context length, while DAnm’s overhead becomes increasingly positive.

Relative accuracy and absolute attended-token change versus vanilla across context-length bins.
Relative accuracy and absolute attended-token change versus vanilla across context-length bins.

5.4. Efficiency results

On a single B200 in bf16, assuming 40% MFU and 70% MBU, the roofline model estimates 192.3 ms for DA versus 269.1 ms for vanilla on Gemma-4-31B, and 237.3 ms versus 306.2 ms on Qwen-3.6-27B. These estimates exclude prefill and model a saturated, optimized, phase-disaggregated serving regime.

The table decomposes modeled decode cost into compute, global-memory, and local-memory terms. DA lowers global KV reads but incurs more compute and fixed local-memory work due to longer generations, yielding the reported modeled totals under B200 roofline assumptions.

Roofline-estimated per-response decode costs for vanilla and Declarative Attention on a B200 accelerator.
Roofline-estimated per-response decode costs for vanilla and Declarative Attention on a B200 accelerator.

6. Analysis

The analysis separates savings due to mode usage from failures to follow the declaration protocol. It identifies global-mode use and invalid focus references as practical limits, particularly for smaller backbones.

6.1. Efficiency impact of DA modes

For Gemma-4-31B, global mode accounts for about 27% of generated tokens overall, while focus and local modes account for the remaining 73%. Focus and local save about 76–99% attention per generated token relative to a vanilla step, although global-mode use rises to about 45% in the longest context bucket.

The figure breaks Gemma-4-31B generations into global, focus, and local modes across context buckets. It identifies focus and local modes as the source of savings while showing that rising global-mode use limits aggregate savings at the longest contexts.

Mode-token shares and per-token attention savings by context length for Gemma-4-31B.
Mode-token shares and per-token attention savings by context length for Gemma-4-31B.

6.2. Protocol adherence

Focus-reference success rises from 58% for Gemma-4-E4B to 99% for Gemma-4-31B, and from 89% for Qwen-3.5-4B to 99% for Qwen-3.6-27B. Focus attempts remain in a narrow range of about 1.4–1.9 per response, indicating that larger models chiefly improve reference validity rather than reduce focus calls.

The figure relates model size to valid focus-reference resolution and focus-call frequency. Increasing focus success rates, especially from 58% to 99% in Gemma, support the conclusion that protocol adherence is a major small-model bottleneck.

Focus-reference success rates and focus attempts per response across model sizes.
Focus-reference success rates and focus attempts per response across model sizes.

7. Conclusion

The paper concludes that model-generated control text can drive reversible KV-cache masking without parameter updates. The authors frame zero-shot results as a lower bound, since the accuracy gap narrows with model capability and absolute savings increase with context length.

8. Discussion, Limitations, and Future Work

DA is limited by zero-shot elicitation, artificial segmentation of static benchmarks, non-thinking-mode evaluation, and costly global navigation steps. Tasks whose evidence is split destructively across segments or whose output length grows with the document can lose accuracy or erase end-to-end savings.

8.1. Relevance of DA in the present and future

The authors argue that DA remains relevant where global attention is a substantial decode-time cost. Their 1M-token roofline calculations estimate global-attention shares of 94% for Kimi-K3 and 56–97% for several indexer-based sparse architectures, but these are architecture-specific modeled values rather than direct measurements.

8.2. Further potential of DA with post-training and in agentic settings

Post-training could reduce DA’s additional decode steps and improve mode-selection policies. Agentic contexts may provide natural segments such as tool outputs and dialogue turns, but these uses are not evaluated; thinking mode is disabled because the tested models did not reliably follow DA syntax inside thinking traces.

8.3. Synergy with modern efficient inference techniques

DA could be combined with lightweight-scan sparse attention for the remaining global phases and with speculative decoding to reduce sequential decoding overhead. The paper argues that block-table masking is compatible with these methods but provides no combination experiments.

8.4. DA as system-2 sparse attention

The authors characterize DA as “system-2 sparse attention,” because the model explicitly reasons about and states its attended region rather than relying solely on latent activation-derived selection. They further propose reversible KV-cache offloading and prefetching based on declared focus spans, but present these as future systems directions.

Appendix

  • Appendix A positions DA against dynamic sparse attention, KV-cache eviction, and attention-steering methods; DA reversibly masks resident blocks rather than permanently evicting them. The appendices also document vLLM integration, roofline derivations, datasets, judge validation, sampling, prompt construction, additional results, and failure cases involving segmentation-sensitive evidence or document-length-proportional outputs.

Brief Thoughts

DA offers a concrete interface between generated reasoning structure and inference-time memory access, and the DAnm ablation shows that masking rather than chunked prompting produces the main attention savings. Its practical usefulness remains conditional on strong protocol adherence, manageable global-mode use, suitable task decomposition, and deployment conditions close to the roofline assumptions.