In-Place Test-Time Training | Summary
07 Apr 2026 | Paper Review Inference-Time Learning Long-Context Language Models Fast WeightsContents
- Summary
- 1 Introduction
- 2 Preliminary: Test-Time Training
- 3 In-Place Test-Time Training
- 4 Experiments
- 5 Related Work
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of In-Place Test-Time Training.
- 2026-04-07 (arXiv)
- Feng, Guhao, Luo, Shengjie, Hua, Kai, Zhang, Ge, He, Di, Huang, Wenhao, Cai, Tianle.
- ByteDance Seed, Peking University
- Paper
- Github
Summary
- In-Place Test-Time Training updates the final projection matrix of existing Transformer MLP blocks in response to context while retaining attention. It provides a checkpoint-compatible fast-weight mechanism rather than replacing attention with an architecture that must be pretrained from scratch.
- The method applies the current weights to each chunk before updating them for subsequent chunks. A Language Modeling-Aligned target derived from token embeddings produces additive updates that can be accumulated through a parallel prefix scan; an induction-style analysis motivates next-token information rather than current-token reconstruction.
- Under matched continual training, Qwen3-4B-Base gains RULER accuracy from 74.3% to 78.7% at 64k, from 74.8% to 77.0% at 128k, and from 41.7% to 43.9% at 256k. From-scratch experiments also show lower long-context perplexity and improved RULER scores, but document-boundary resets limit the demonstrated adaptation to within-document memory rather than persistent learning across an unbounded stream.
1 Introduction
The paper identifies three obstacles to using Test-Time Training in pretrained LLMs: specialized architectures that cannot readily inherit checkpoints, sequential updates that underutilize accelerators, and reconstruction objectives that do not explicitly support next-token prediction. It addresses them by reusing MLP parameters, updating in large chunks, and constructing predictively oriented targets.
- Attention remains responsible for token mixing, allowing the adaptive MLP to complement rather than replace pretrained attention.
- Language modeling and long-context benchmarks serve as proxies for long-horizon adaptation; the experiments do not directly evaluate the evolving agent tasks motivating the introduction.
- The released code URL is https://github.com/ByteDance-Seed/In-Place-TTT.
2 Preliminary: Test-Time Training
Test-Time Training treats selected parameters as an evolving contextual memory. The preliminary section separates the update-and-retrieval mechanism from the requirements for integrating it into pretrained autoregressive language models.
The TTT mechanism.
In the conventional formulation, token representations produce queries, keys, and values. Fast weights are first updated to associate a key with its value through a self-supervised loss, and the updated network then processes the query.
- The update is Wi ← Wi−1 − η∇W L(fWi−1(ki), vi), followed by oi = fWi(qi).
- Losses, optimizers, and memory parameterizations determine what information this state stores, retrieves, and forgets.
- In-Place TTT reverses this ordering at chunk granularity: apply first, then update for future chunks.
Desiderata for TTT within the LLM ecosystem.
The paper’s desiderata are Architectural Compatibility, Computational Efficiency, and a Tailored Learning Objective for Language Modeling. Compatibility means warm-starting from a pretrained checkpoint; efficiency requires moving beyond sequential per-token updates; objective alignment requires targets chosen for their predictive relevance.
- These requirements motivate preserving the Transformer backbone while changing how existing MLP weights operate during inference.
- The critique of reconstruction concerns predictive usefulness, not whether reconstruction can store contextual information.
3 In-Place Test-Time Training
In-Place Test-Time Training combines an adaptable MLP down projection, chunk-wise updates, an LM-Aligned Objective, and a context-parallel implementation. The framework preserves attention and pretrained MLP initialization while adding components that generate online update targets.
3.1 Overall Framework
Repurposing MLP Blocks for In-Place Adaptation makes Wdown the fast-weight state of a gated MLP. Wup and Wgate remain slow weights during online adaptation, and the intermediate activations provide the keys used to modify the final projection.
- The gated MLP computes O = (ϕ(HWgateᵀ) ⊙ (HWupᵀ))Wdownᵀ.
- Using an existing projection avoids replacing attention or introducing a new primary token-mixing layer.
- The target-generating Conv1D and Wtarget are additional learned components, so checkpoint compatibility does not mean that the method introduces no parameters.
Efficient Adaptation with Chunk-Wise Updates partitions activations and targets into non-overlapping chunks of size C. Each chunk is processed with the fast-weight state obtained from earlier chunks, after which a gradient step updates that state.
- The apply operation is O[i] = Z[i](Wdown(i))ᵀ; the update uses Z[i] and V[i] to construct Wdown(i+1).
- Because attention handles within-chunk interactions, the adaptive MLP need not use the small chunks required when TTT is the sole token mixer.
- The reported ablations favor C = 512 and C = 1024, with C = 1024 offering better efficiency.
3.2 LM-Aligned Objective
The LM-Aligned Objective defines V̂ = Conv1D(X0)Wtarget, where X0 contains token embeddings. The intended target incorporates localized future-token information, with a next-token embedding as a special case, rather than reconstructing the current token.
- Using L(·, ·) = −⟨·, ·⟩F gives the additive update Wdown(i) = Wdown(i−1) + ηV̂[i]ᵀZ[i] in equation (1).
- Because this gradient does not depend on the current fast-weight state, update deltas can be computed independently and accumulated associatively.
- The online objective is a representation-level similarity loss, not a cross-entropy step through the entire language model.
3.3 Theoretical Analysis
Setup. The theoretical analysis considers a single enhanced Transformer block in an induction task: a key-value pair (k*, v*) occurs earlier, and a later occurrence of k* should predict v*. It compares current-token and next-token embedding targets; the matching key must contribute an update from an earlier chunk, and the reconstruction comparison assumes k* and v* are distinct.
- The contextual update changes the query-position logit by Δℓn[w] = EwᵀΔWdownZn.
- Assumptions. Distinct token embeddings are approximately orthogonal, |ewiᵀewj| ≤ ϵ, and their squared norms satisfy ‖ewi‖² ≥ cnorm² > 0.
- The matching key and query activations satisfy E[zt*ᵀzn] = calign > 0, while unrelated positions have zero expected contribution, E[(VtZtᵀ)Zn] = 0.
Theorem 1 states that the LM-Aligned target raises the correct token’s expected logit by at least λlr · cnorm² · calign, while the absolute expected change for each other token is at most λlr · ϵ · calign. With reconstruction, the correct token’s absolute expected logit change is bounded by the latter quantity.
- The result explains why storing the successor of a matching key helps an induction-style prediction.
- Its scope is the stated single-block setting and representation assumptions; it does not prove general gains for a multilayer LLM with learned convolutional targets.
- The practical target extends the next-token case to a learned localized combination of future tokens.
3.4 Implementation Details
Efficient Implementation with Context Parallelism computes each chunk’s activations and update delta in parallel, accumulates earlier deltas with an exclusive prefix sum, and applies the resulting effective weights in parallel. Figure 1 illustrates the corresponding apply-then-update data flow.
- The prefix state is ΔSi = ∑j=1…i−1 ΔWj, and the effective projection is Wdown(0) + ηΔSi.
- Causality and Boundary Handling. The paper specifies chunk-local target generation with causal padding and resets fast weights to their pretrained state at document boundaries.
- Discussion. The authors describe the scan as equivalent to sequential causal updates and leave other losses and optimizers for future work.
The diagram separates unchanged attention from the adaptive MLP down projection. Token embeddings feed the target generator, while gated MLP activations provide update keys; each chunk's output uses the current projection before its delta changes the state for later chunks. This ordering prevents a chunk from using its own target-dependent update.
4 Experiments
The experiments test whether In-Place TTT improves pretrained LLMs, how it compares with prior methods when trained from scratch, and which design choices affect performance. Evaluation covers continual training at 4B–14B scales, from-scratch models at 500M, 1.5B, and 4B, and ablations with a 1.7B model.
- The principal long-context measures are RULER accuracy and Sliding Window Perplexity.
- Common sense reasoning evaluations provide an additional check on the capabilities of the 4B models.
4.1 In-Place TTT as a Drop-in Enhancement for Pre-trained LLMs
Training and Evaluation. Qwen3-4B-Base and its In-Place TTT counterpart receive the same continual-training curriculum: ∼20B tokens at 32k context, followed by ∼15B tokens at 128k context. Stage 2 uses YaRN, and OpenCompass evaluates RULER from 4k to 256k, with 256k testing extrapolation beyond the training length.
- Both stages use AdamW with learning rate 5e-6 and weight decay 0.1; the exact sequence lengths are 32,768 and 131,072.
- The method is a checkpoint-compatible enhancement that still requires substantial continual training, not a demonstrated zero-training modification.
- Reported Qwen3-4B evaluations clip each update delta’s Frobenius norm to τ = 1e-5 for numerical stability.
Results and Discussion. Table 1 shows the largest absolute gain at 64k, where accuracy rises from 74.3% to 78.7%; the 128k and 256k scores rise from 74.8% to 77.0% and from 41.7% to 43.9%. The enhancement improves six of the seven tested lengths, but at 4k it scores 96.1% against the baseline’s 96.6%.
- The longer-context gains support useful contextual adaptation under the matched training curriculum.
- The advantage does not increase monotonically with context length, despite the paper’s description of a widening trend.
- At 256k, both models remain substantially below their 128k performance, demonstrating a relative extrapolation gain rather than reliable unlimited-context processing.
Extension to More Models. Table 2 reports gains at every tested length for LLaMA-3.1-8B and Qwen3-14B-Base, including 81.6% to 83.7% and 67.9% to 70.6% at 64k. For Qwen3-14B with YaRN at 64k, accuracy increases from 81.3% to 82.5%.
- These comparisons extend the evidence beyond one checkpoint family and one parameter scale.
- Although the main text describes the same continual-training protocol, Table 7 specifies ∼20B tokens at sequence length 32,768 for each additional model, rather than the full Qwen3-4B two-stage curriculum.
- The retained gain with YaRN supports compatibility with positional extrapolation techniques.
4.2 Pre-training from Scratch: A Comparative Analysis
Experimental Setup. At 500M and 1.5B, the paper compares In-Place TTT with Sliding-Window Attention (SWA), Gated Linear Attention (GLA), DeltaNet, and Large Chunk Test-Time Training (LaCT). In-Place TTT and LaCT both use an SWA backbone, and all models train at sequence length 32,768.
- The 500M models train for 20B tokens with a 2,048-token attention window; the 1.5B models train for 60B tokens with a 4,096-token window.
- AdamW learning rates are 5e-4 and 3e-4, with batch sizes of 2M and 4M tokens respectively; both use weight decay 0.1, gradient clipping 1.0, and 1024 warmup steps.
- At 4B, full-attention and SWA backbones are compared with their enhanced counterparts after 120B training tokens at sequence length 8,192.
Evaluation. Sliding Window Perplexity evaluates a fixed final token block as the preceding context grows, using a validation set comprising Pile and Proof-Pile-2. Figure 2 displays the Pile results: In-Place TTT has the lowest plotted perplexity at both model scales and continues benefiting from context up to 32k.
- SWA largely plateaus after its local attention window, whereas the adaptive models can use information beyond that window.
- The continued decline supports long-context utilization, but the plot does not isolate which stored information accounts for the gain.
- Exact perplexity values are not tabulated, so the comparison is reported as the visible ordering and trend.
At both 500M and 1.5B scales, the In-Place TTT curve stays below SWA, GLA, DeltaNet, and LaCT as context grows. Its perplexity continues decreasing through 32k, whereas SWA flattens once additional context lies outside its attention window. The figure supports longer-range context utilization rather than only improved short-context modeling.
Results and Discussion. Table 3 shows that the 4B full-attention model’s RULER-16k score increases from 6.58 to 19.99, while the SWA model’s RULER-8k score increases from 9.91 to 26.80. Most common sense reasoning scores improve, but full-attention ARC-C decreases from 33.19 to 32.34 and SWA PIQA decreases from 72.58 to 72.03.
- The long-context gains appear with both full attention and SWA, supporting the claim that the mechanism complements attention.
- The 16k RULER evaluation exceeds the 8k training length for these models, but absolute scores remain low.
- The mixed common sense results support improvement on most tasks, not uniform improvement across downstream capabilities.
In-Place TTT improves every RULER entry for both attention backbones, including full-attention RULER-16k from 6.58 to 19.99 and SWA RULER-8k from 9.91 to 26.80. Common sense gains are smaller and not universal: full-attention ARC-C and SWA PIQA decline. The strongest evidence is for long-context performance, not uniform downstream improvement.
4.3 Ablation Studies: On the Impact of Key Design Choices
Figure 3 tests fast-weight state size, chunk size, and target construction on RULER with a 1.7B model. Increasing the number of TTT-enabled layers improves the plotted scores, while C = 512 and C = 1024 outperform C = 256 and C = 2048.
- Impact of State Size. The plotted settings are 4×, 1×, and 0.5×; larger adaptive state is obtained by changing the number of enabled layers.
- Impact of Chunk Size. The preferred intermediate sizes reflect a trade-off between adaptation frequency and parallel efficiency, rather than a general rule that smaller chunks are better.
- Impact of LM-Aligned Objective. Removing Conv1D particularly harms long-context scores, while removing Wtarget sharply reduces short-context scores; the full target construction provides the strongest overall result.
The state-size panel favors more TTT-enabled layers, and the chunk-size panel favors intermediate update intervals rather than either extreme. Removing convolution is especially damaging at longer contexts, while removing projection strongly harms shorter-context scores. The ablations support the combined target design in this setting, not a general performance guarantee.
Efficiency Impact of In-Place TTT. Figure 4 reports prefill throughput and peak memory for Qwen3-4B-based continual-trained checkpoints on Nvidia H800 GPUs, at batch size 1 and context lengths 8k, 32k, and 128k. The enhanced models show modest throughput reductions and small peak-memory increases compared with their corresponding baselines.
- The efficiency-only SWA setting manually changes attention to a 1024-token sliding window.
- With full attention, throughput falls substantially and memory increases as context grows in both baseline and enhanced models; In-Place TTT does not remove attention’s long-context cost.
- The evidence concerns prefill under this hardware and batch-size configuration, not decoding latency or general serving throughput.
The enhanced models incur a visible but modest prefill-throughput penalty and little additional peak memory in the tested configuration. Full-attention costs still grow strongly with sequence length, while SWA keeps throughput comparatively stable. These measurements support low incremental adaptation overhead, not removal of the backbone's attention cost.
5 Related Work
The related-work discussion covers online fast-weight adaptation, efficient long-context architectures, and contextual memory design. In-Place TTT adapts pretrained MLP projections while retaining the existing attention backbone.
Test-Time Training (TTT).
Prior Test-Time Training work explores self-supervised adaptation, optimizer design, and expressive neural memories across vision, language, video, and audio. Chunk-wise formulations address sequential execution, but using TTT as the primary token mixer couples chunk size more tightly to local information exchange.
- In-Place TTT retains attention for local mixing while assigning MLP fast weights a complementary memory role.
- Its contribution is not chunk-wise updating alone, but its combination with checkpoint-compatible MLP adaptation and an LM-Aligned target.
Efficient Long-Context Architectures.
Sparse attention, linear attention, gated recurrent formulations, and State-Space Models reduce the cost of processing long sequences. The paper treats these architectures as complementary to online adaptation and proposes extending its MLP-based mechanism to other efficient backbones.
- The experiments directly demonstrate compatibility with full attention and SWA.
- Integration with the broader set of efficient backbones discussed here remains future work.
Memory Design and Augmentation.
The memory discussion distinguishes persistent, task-agnostic external knowledge from transient contextual information. In-Place TTT belongs to the latter category: model parameters temporarily function as a state updated as a document is processed.
- Unlike an RNN activation vector, the state is a projection matrix within the model.
- Resetting that matrix at document boundaries separates contextual adaptation from durable knowledge acquisition across documents.
6 Conclusion
The conclusion attributes the results to repurposing existing MLP projections, using scalable chunk-wise updates, and aligning update targets with language modeling. The experiments support checkpoint-compatible contextual adaptation and competitive from-scratch performance; they do not directly test persistent continual learning.
Appendix
- A Proof of theorem 1 expands the logit change as a sum of embedding-target similarities multiplied by key-query activation similarities. The zero-contribution assumption removes unrelated positions, and approximate embedding orthogonality yields the bounds for next-token and reconstruction targets, with distinct key and value tokens required for the reconstruction bound. B Context Parallel Algorithm for In-Place TTT provides Algorithm 1, whose effective weights use updates from chunks < i and reset at document boundaries.
- C Experiment Details and C.1 Details of Datasets describe an internally collected from-scratch pretraining mixture of English and Chinese text, multilingual material, code, mathematics, and knowledge- or reasoning-dense data. The continual-pretraining mixture combines similar short documents with books, repository-level code, and synthetic retrieval-augmented and long-context-QA constructions, organized for 32k and 128k training. The main text additionally identifies TogetherAI data for the 500M and 1.5B comparisons, but the appendix does not fully reconcile that description with its dataset summary.
- C.2 Details of Training and Evaluation specifies Nvidia H800 GPUs, lm-evaluation-harness for common sense reasoning, and OpenCompass for long-context evaluation. Tables 4–7 provide the training schedules; Table 5 lists AdamW, learning rate 3e-4, batch size 8M tokens, weight decay 0.1, gradient clipping 1.0, 1.6B warm-up tokens, sequence length 8,192, and 120B training tokens for the 1.7B and 4B pretraining settings. Qwen3-4B evaluation rescales any update delta exceeding Frobenius norm τ = 1e-5 to that threshold.
- C.3 Details of Model Configuration specifies decoder-only Transformers with SwiGLU and RoPE. Table 8 lists 24 layers for both smaller models, hidden sizes 1024 and 2048, FFN sizes 3072 and 6144, 8 and 16 attention heads, vocabulary size 32,000, and RoPE base 1e6. The 1.7B and 4B backbones follow Qwen3-1.7B-Base and Qwen3-4B-Base, and the standard integration enables In-Place TTT in every sixth layer.
- Initialization of In-Place TTT Modules uses a zero-initialized depthwise Conv1D with kernel size 5, causal padding, and no bias. Wtarget starts as a sparse diagonal matrix with diagonal entries drawn from N(0, σ²), making the initial update zero and preserving pretrained behavior before the target generator learns. D Usage of LLMs reports assistance with grammar and readability, with the authors retaining responsibility for the manuscript.
Brief Thoughts
The matched Qwen3-4B comparison is the clearest evidence for the enhancement: both variants receive the same continual-training curriculum, and the adaptive version improves most RULER lengths. However, “drop-in” describes architectural compatibility rather than negligible training cost: the main experiment uses ∼35B continual-training tokens. The reported 256k improvement leaves accuracy at 43.9%, limiting the strength of the extrapolation claim. No uncertainty intervals or repeated-run variability are reported for these comparisons. The internally collected data descriptions do not provide a complete reproducible dataset specification.
The analysis motivates predictive targets, but the implementation description requires clarification: section 3.2 asks Conv1D targets to contain future-token information, whereas section 3.4 and Algorithm 1 specify causal padding. Future information confined to a completed chunk can be compatible with updates applied only to later chunks, but the manuscript does not resolve the exact convolution indexing used to implement both descriptions. The prefix scan must be exclusive, as specified by the main-text sum and Algorithm 1’s chunks < i comment, even though the pseudocode names the operation Cumsum. Document-boundary resets and the absence of cross-document retention experiments restrict the demonstrated mechanism to contextual memory. Efficiency conclusions are limited to batch-size-1 H800 prefill measurements; decoding performance and stability over unbounded streams are not established.