Gorio Tech Blog search

$\delta$-mem: Efficient Online Memory for Large Language Models | Summary

|

Contents

This article explains the key points of $\delta$-mem: Efficient Online Memory for Large Language Models.

  • 2026-05-12 (arXiv), 40th Conference on Neural Information Processing Systems (NeurIPS 2026)
  • Lei, Jingdi, Zhang, Di, Li, Junxian, Wang, Weida, Fan, Kaixuan, Liu, Xiang, Liu, Qihan, Ma, Xiaoteng, Chen, Baian, Poria, Soujanya.
  • Nanyang Technological University, Fudan University, Shanghai Jiao Tong University, The Chinese University of Hong Kong, The Hong Kong University of Science and Technology (Guangzhou), Mind Lab
  • Paper
  • Github

Read this article in Korean


Summary

  • δ-mem augments a frozen full-attention Transformer with an online state of associative memory (OSAM). Instead of replaying stored history as text, it reads a compact recurrent matrix, converts the readout into query-side and output-side attention corrections, and updates the matrix through a gated delta rule.
  • With rank r = 8, the single-state variants use an 8 × 8 memory matrix at each insertion layer and 4.87M trainable parameters. On Qwen3-4B-Instruct, Token-State Write (TSW) increases the five-benchmark average from 46.79 to 51.66, while Multi-State Write (MSW) achieves the best MemoryAgentBench and LoCoMo averages of 38.85 and 49.12.
  • The experiments show useful compressed-history retention and aggregate improvements across three backbones, but neither exact context recovery nor improvements on every task. TSW decodes more slowly than the unmodified backbone, and the no-context experiment reports inconsistent LoCoMo values in the prose and Figure 2.

1 Introduction

Long-term assistants and agents must accumulate, revise, and reuse information across interactions. The introduction argues that larger context windows are an incomplete solution because full attention becomes more expensive and additional tokens do not guarantee effective context utilization.

  • The paper separates memory mechanisms into memory state, which determines how history is stored, and memory steering, which determines how that history affects the backbone.
  • Textual memory is constrained by retrieval quality and context budgets; outside-channel memory adds retrieval and integration machinery; parametric memory generally lacks continuously evolving inference-time state.
  • δ-mem maintains a compact, persistent state whose readout directly influences attention while keeping pretrained backbone parameters frozen. The released code URL is https://github.com/declare-lab/delta-Mem.

2 Preliminaries

The preliminary formulation treats the memory matrix as an online key–value regressor: its prediction for a key is the previous state multiplied by that key. An SGD step on squared prediction error yields the delta update S_t = S_{t−1} + β_t(v_t − S_{t−1}k_t)k_tᵀ; a retention gate adds controlled forgetting.

  • The residual term corrects associations that the existing state predicts incorrectly, rather than simply accumulating every new key–value outer product.
  • The gated preliminary update is S_t = λ_tS_{t−1} + β_t(v_t − S_{t−1}k_t)k_tᵀ. The method subsequently implements dimension-wise gates.

3 δ-mem

δ-mem follows a read–steer–write sequence at each position: query the previous memory state, use the resulting signal to correct attention, and then incorporate the current information. Figure 1 shows the memory branch alongside the frozen Transformer and distinguishes token-level, segment-level, and multiple-state writing.

  • Reading the old state before writing the current position makes the steering signal depend on previously accumulated information.
  • The backbone retains standard full attention; δ-mem augments it rather than replacing its token-mixing architecture.

The diagram separates the frozen attention pathway from the trainable memory interface. A previous-state readout affects query formation and the attention output, while current keys, values, and gates update the state afterward. The right-hand panels distinguish update frequency from multiple-state organization.

δ-mem architecture with associative-state reading, low-rank attention steering, delta-rule writing, and three writing strategies.
δ-mem architecture with associative-state reading, low-rank attention steering, delta-rule writing, and three writing strategies.

3.1 Memory Projections

Learned projections map the current hidden state into memory queries, keys, and values of dimension r. Queries and keys use tanh followed by L2 normalization, while values use a linear projection; normalization is intended to limit scale drift during recurrence.

  • The write gate is (\beta_t=\sigma(W_\beta x_t+b)), and retention is tied to it through (\lambda_t=1-\beta_t).
  • Both gates are r-dimensional vectors, allowing different memory rows to retain old information or write new residuals at different rates.

3.2 Reading from Online State of Associative Memory

The read operation is (r_t=S_{t-1}q_t^m), producing a continuous associative-memory signal from the previous state. Its cost is independent of accumulated history length because the state dimensions remain fixed.

  • Unlike standard attention over explicit keys, this operation does not search all historical token representations.
  • The readout supplies a history-dependent steering signal rather than retrieved text or additional context tokens.

3.3 Steering Attention through Low-Rank Corrections

Two linear mappings transform the memory readout into corrections Δq_t and Δo_t. The query correction is added before attention, and the output correction is added afterward, both scaled by α/r; the backbone keys and values remain unchanged in the default configuration.

  • The corrected query is q̃_t = W_Qx_t + (α/r)Δq_t, and the corrected output is ỹ_t = Attn(q̃_t, K_≤t, V_≤t) + (α/r)Δo_t.
  • Although the correction projections are fixed after training, their inputs depend on the evolving memory state. This distinguishes the mechanism from a static low-rank adapter.

3.4 Writing into Online State of Associative Memory

Writing uses the dimension-wise gated delta rule S_t = Diag(λ_t)S_{t−1} + Diag(β_t)(v_t^m − S_{t−1}k_t^m)(k_t^m)ᵀ. This combines retention of the previous state with a residual correction along the current key direction.

  • Expanding the update separates three operations: retain old state, subtract its prediction component along the current key, and write the new value along that direction.
  • Each row has an independent gate, providing a mechanism for balancing stable associations against newly arriving information.

3.5 Writing Granularity of Online State

The paper evaluates three writing strategies that change update granularity or memory organization. TSW updates after every token, Sequence-State Write (SSW) updates once per message using its mean hidden representation, and MSW maintains parallel sub-states whose readouts are concatenated.

  • TSW retains fine-grained changes but can expose the state to repeated expressions, formatting tokens, and short-term noise.
  • SSW reduces redundant writes and smooths updates, at the cost of averaging away token-level detail.
  • MSW is intended to reduce interference between different kinds of information. The implementation uses 4 sub-states, each with an 8 × 8 matrix.

3.6 Training Objective

Training uses supervised fine-tuning with autoregressive cross-entropy over response tokens. Historical context is first written into a state S_C, but is not replayed as explicit backbone input during prediction; the frozen backbone receives the query and response sequence while δ-mem supplies state-conditioned steering.

  • The objective is L_SFT = −Σ_j log p_{ϕ,θ}(y_j | Q, y_<j, S_C), where ϕ denotes frozen backbone parameters and θ denotes trainable δ-mem parameters.
  • This training setup requires the memory pathway to carry historical information into response generation.

4 Experiments

The experiments assess memory-heavy tasks and general capabilities. They compare memory mechanisms on a shared Qwen3-4B-Instruct backbone and evaluate the design on Qwen3-8B and SmolLM3-3B.

4.1 Experimental Setup

Evaluation covers HotpotQA, GPQA-Diamond, IFEval, LoCoMo, and MemoryAgentBench; the LoCoMo adversarial question category is excluded. Baselines include BM25 RAG, LLMLingua-2, MemoryBank, Context2LoRA, MemGen, and MLP Memory.

  • The reported training setup uses one epoch on the shortest 2,219-sample QASPER split. Its maximum sequence length is 8,269 tokens; backbone training inputs are limited to 512 tokens, with an 8,192-token memory write budget.
  • Default settings are r = 8, α = 16, query/output corrections, and 4 states for MSW. Training uses 8 × A800 GPUs, bfloat16, DeepSpeed ZeRO-2, and fused AdamW.
  • Metrics are prompt-level strict accuracy for IFEval, EM/F1 for HotpotQA, accuracy for GPQA, F1 for LoCoMo, and a sample-weighted aggregate of dataset-specific metrics for MemoryAgentBench.

4.2 Main Results across Memory Mechanisms

Table 1 reports the highest overall average for TSW: 51.66 versus 46.79 for the frozen backbone and 44.90 for Context2LoRA, gains of 4.87 and 6.76 points. SSW reaches 51.44 and MSW reaches 50.74; HotpotQA contributes EM rather than F1 to this final average.

  • MSW raises MemoryAgentBench from 29.54 to 38.85 and LoCoMo from 40.79 to 49.12. SSW gives the largest Test-time Learning (TTL) score, increasing it from 26.14 to 50.50.
  • TSW improves HotpotQA EM/F1 from 42.35/56.00 to 49.41/63.66 and IFEval from 81.89 to 82.99.
  • The gains are not uniform: all three variants score below the backbone’s 47.08 on MemoryAgentBench Long Range Understanding (LRU), and MSW lowers GPQA-Diamond from 39.39 to 37.37. The authors’ explanations involving baseline retrieval noise or overfitting are not independently tested.

TSW leads the overall average at 51.66, but MSW leads the two memory-heavy benchmark averages and SSW leads the TTL subtask. Effects are nonuniform: every δ-mem variant lowers LRU relative to the backbone, and MSW lowers GPQA-Diamond accuracy. Some baseline entries are unreported, and the caption does not explain how missing entries are handled in the final averages.

Benchmark results for memory mechanisms built on Qwen3-4B-Instruct, including general capabilities and memory-heavy tasks.
Benchmark results for memory mechanisms built on Qwen3-4B-Instruct, including general capabilities and memory-heavy tasks.

4.3 Main Results Across Different Backbone Models

Table 2 shows higher aggregate scores for δ-mem on all three backbones. The best variant raises Qwen3-4B-Instruct from 46.79 to 51.66, Qwen3-8B from 47.20 to 50.86, and SmolLM3-3B from 26.08 to 36.96.

  • TSW leads on Qwen3-4B-Instruct, SSW leads on Qwen3-8B, and MSW leads on SmolLM3-3B. No writing strategy dominates across the tested backbones.
  • SmolLM3-3B shows especially large HotpotQA gains: MSW increases EM from 1.67 to 31.61 and F1 from 14.40 to 46.77.
  • The authors attribute these differences to smoother segment-level updates for a stronger backbone and reduced state interference for a smaller one. The comparison does not isolate model capacity as the cause.

Aggregate improvements appear on all three backbones, but the strongest writing strategy changes: TSW for Qwen3-4B-Instruct, SSW for Qwen3-8B, and MSW for SmolLM3-3B. SmolLM3-3B shows the largest aggregate increase, from 26.08 to 36.96. These comparisons do not establish a uniform scaling relationship between model size and memory benefit.

General and memory-heavy benchmark results across Qwen3-4B-Instruct, Qwen3-8B, and SmolLM3-3B.
General and memory-heavy benchmark results across Qwen3-4B-Instruct, Qwen3-8B, and SmolLM3-3B.

5 Ablative Study

The ablations examine information retention without explicit history, the attention branches receiving corrections, and module insertion depth. These tests use Qwen3-4B-Instruct and focus on HotpotQA and LoCoMo.

5.1 Context Recovery

With historical context removed, Figure 2 shows HotpotQA overall EM increasing from 0.08 to 6.48 and F1 from 8.27 to 15.20. These gains support partial retention of useful evidence, but do not establish equivalence to full-context inference.

  • Bridge EM/F1 increases from 0.08/6.25 to 3.97/11.05; Comparison EM/F1 increases from 0.07/16.31 to 16.48/31.70.
  • For LoCoMo, the figure displays an average increase from 6.89 to 9.62, whereas the accompanying prose reports 3.49 to 8.05. The PDF does not reconcile these values.
  • The figure also shows improvements across Multi, Temp, Open, and Single categories, with the largest displayed LoCoMo category gain on Open: 11.97 to 18.00.

The stacked bars show useful but incomplete recovery after explicit historical context is removed. HotpotQA overall EM increases from 0.08 to 6.48, while the plotted LoCoMo average increases from 6.89 to 9.62. The latter conflicts with the prose values of 3.49 and 8.05, leaving the LoCoMo recovery estimate unresolved.

No-context baseline scores and gains from compressed-memory state on HotpotQA and LoCoMo.
No-context baseline scores and gains from compressed-memory state on HotpotQA and LoCoMo.

5.2 Heads Ablation

Table 3 compares corrections to query, key, value, and output branches and their combinations. The default qo configuration reaches an average of 47.97, close to the full qkvo configuration’s 48.05; the authors choose qo to avoid extra correction branches for a 0.08-point aggregate gain.

  • Among single-branch variants, output correction performs best at 47.05, compared with 44.51 for q, 42.19 for k, and 44.24 for v.
  • qkvo obtains the best HotpotQA overall EM/F1, 49.94/65.01, while qo obtains the best LoCoMo average, 46.53, versus 46.15 for qkvo.
  • This is an ablation of correction placement within attention, not a comparison of the number of conventional attention heads.

Output-only correction is the strongest single-branch option. Correcting query and output together nearly matches the full qkvo configuration in aggregate score, 47.97 versus 48.05, and achieves the higher LoCoMo average. This supports qo as a lower-overhead choice, not the highest-scoring configuration on every metric.

Ablation of memory-induced corrections to query, key, value, and output branches.
Ablation of memory-induced corrections to query, key, value, and output branches.

5.3 Insertion Depth Ablation

Table 4 compares insertion in the front 12, middle 12, back 12, or all backbone layers. All-layer insertion achieves the best average of 47.97, followed by middle-layer insertion at 46.66, front-layer insertion at 44.39, and back-layer insertion at 44.06.

  • All-layer insertion gives HotpotQA EM/F1 of 49.41/63.66 and a LoCoMo average of 46.53.
  • Middle-layer insertion is the strongest partial-depth option and achieves the best LoCoMo Multi score, 44.00.
  • The authors attribute the pattern to intermediate semantic representations and propagation of corrections across depth. Because all-layer insertion also adds more modules, this is not a parameter-matched isolation of placement.

All-layer insertion produces the highest aggregate score and the strongest overall HotpotQA and LoCoMo results. Middle-layer insertion performs best among the three 12-layer alternatives. Because all-layer insertion uses more modules, that comparison combines placement and adaptation-budget effects.

Performance with δ-mem inserted in front, middle, back, or all backbone layers.
Performance with δ-mem inserted in front, middle, back, or all backbone layers.

The related-work section compares δ-mem with textual stores, external latent-memory banks, parametric adaptation and editing, and hybrid attention or recurrent architectures. Its distinguishing combination is a frozen full-attention backbone, interaction-persistent associative state, gated delta writing, and state-conditioned attention corrections.

  • Compared with textual approaches, δ-mem avoids reinserting memory through token space; compared with Memorizing Transformers and LongMem, it does not retrieve an external bank through a separate reader pathway.
  • Its low-rank interface resembles LoRA, but runtime steering depends on changing state rather than a static weight update.
  • Infini-attention also uses compressed memory and considers delta-rule updates, while Titans updates neural-memory parameters at test time. δ-mem’s contribution is the specific persistent-state and query/output-correction interface, not the invention of associative recurrence or delta learning.

7 Conclusion

The conclusion argues that compact online state can provide useful memory to frozen Transformers without full fine-tuning, architectural replacement, or explicit history replay. The benchmark averages and no-context recovery results support useful memory augmentation for the tested models, with approximate retention and decoding overhead as qualifications.

  • Figures 3 and 4 characterize efficiency and parameter overhead: TSW adds little observed GPU memory but reduces decoding throughput relative to Vanilla and Context2LoRA, while the trainable module counts remain small relative to the backbone.
  • Table 5 documents the heterogeneous dataset metrics behind MemoryAgentBench, which matter when interpreting its aggregate score.

Across six prompt/decode-length combinations, TSW's memory bars closely track Vanilla and Context2LoRA, but its decoding throughput is lower. The figure supports low additional GPU memory usage, not inference-speed parity or an efficiency claim for all writing variants. Its memory axis does not specify a physical unit.

Decoding throughput and memory usage for prompt lengths of 4k, 16k, and 32k with decoding lengths of 64 and 256.
Decoding throughput and memory usage for prompt lengths of 4k, 16k, and 32k with decoding lengths of 64 and 256.

MemoryAgentBench combines Accurate Retrieval, Test-time Learning, Long Range Understanding, and Selective Forgetting. Most datasets use accuracy, but Movie Recommendation uses Recall@5 and ∞Bench-Sum uses F1-Score. Its sample-weighted final score therefore summarizes different measurement types rather than a single common metric.

MemoryAgentBench evaluation categories, constituent datasets, and dataset-specific metrics.
MemoryAgentBench evaluation categories, constituent datasets, and dataset-specific metrics.

Appendix

  • A Limitations: The compact state approximates explicit context, and the authors identify closer alignment with exact full-context inference as future work. They also acknowledge additional decoding operations and defer system-level optimization; low GPU memory overhead does not eliminate the measured throughput cost.
  • B Broader Impacts: Persistent compressed memory may help assistants and agents operate under limited computational budgets, but introduces risks of sensitive-information retention, stale or incorrect memories, and user profiling. The paper recommends consent, transparent controls, deletion mechanisms, and filtering or auditing, without experimentally evaluating those safeguards.
  • C Inference Efficiency and Memory Use: Figure 3 evaluates TSW at prompt lengths of 4k, 16k, and 32k, each with decoding lengths of 64 and 256. Its displayed GPU memory usage is nearly the same as Vanilla and Context2LoRA, while decoding is slower than both and faster than MemGen; the memory axis specifies no physical unit.
  • D Implementation Details: Training uses a peak learning rate of 2 × 10⁻⁴ with cosine decay, a warmup ratio of 0.1, per-device batch size 1, 4 gradient accumulation steps, effective global batch size 32, and random seed 42. Evaluation follows official prompts and decoding settings; Table 5 assigns accuracy to most MemoryAgentBench datasets, Recall@5 to Movie Recommendation, and F1-Score to ∞Bench-Sum, then aggregates by sample count.
  • E Parameter Overhead: Figure 4 reports 4.87M trainable parameters for SSW and TSW, each 0.12% of the backbone, and 19.47M for MSW, or 0.48%. Context2LoRA uses 5.90M (0.15%), MemGen 46.20M (1.13%), and MLP Memory 3078.00M (76.40%); these learned-module counts are distinct from the runtime associative-state matrices.

Brief Thoughts

The main contribution is the separation of trained memory interfaces from dynamically updated memory content. The qo ablation, cross-backbone aggregate improvements, and low parameter counts show that a narrow attention interface can exploit recurrent state while keeping backbone parameters frozen.

The evidence is less complete for faithful long-term storage and end-to-end efficiency. Partial no-context recovery, lower LRU scores on the main backbone, slower TSW decoding, a single reported training seed, and inconsistent LoCoMo recovery numbers limit the conclusion to improved tested aggregates—not lossless history compression or uniformly better memory behavior.