Recursive Language Models | Summary
31 Dec 2025 | Paper Review Recursive Language Models Long-Context Language Models Test-Time Scaling Supervised Fine-TuningContents
- Summary
- 1 Introduction
- 2 Recursive Language Models
- 3 Scaling Long Context Tasks
- 4 Results and Discussion
- 5 Analyses of RLM Trajectories
- 6 Related Works
- 7 Limitations and Future Work
- 8 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Recursive Language Models.
- 2025-12-31 (arXiv), Preprint
- Zhang, Alex L., Kraska, Tim, Khattab, Omar.
- MIT CSAIL
- Paper
Summary
- Recursive Language Models proposes an inference framework that stores the user prompt in an external programming environment rather than placing it directly in the underlying LLM’s context window. The model writes code to inspect and transform the prompt, launch programmatic model calls, and assemble intermediate results and final outputs in persistent variables.
- Experiments distinguish sparse retrieval from dense aggregation and pair enumeration, evaluating GPT-5 and Qwen3-Coder-480B-A35B across four main long-context benchmarks. On BrowseComp-Plus with 6M–11M input tokens, RLM(GPT-5, depth=1) achieves 91.3% at an average API cost of $0.99; on 32K-token OOLONG-Pairs, it achieves 58.0% F1 versus 0.1% for direct GPT-5.
- External context storage alone often improves performance, while sub-calls provide additional gains on dense tasks. Greater recursion depth is not uniformly better, and the sequential main implementation produces long-tailed latency and cost.
- Small-scale post-training targets the root model’s ability to manipulate the environment and delegate work, rather than training every leaf computation. RLM-Qwen3-8B improves over untrained Qwen3-8B in the RLM scaffold by a reported median of 28.3% across the four evaluation tasks, while a separate reinforcement-learning experiment demonstrates length and needle-count generalization on MRCRv2.
1 Introduction
The paper addresses both hard context-window limits and context rot within those limits. Figure 1 shows that GPT-5 retains strong single-needle retrieval performance while degrading earlier on OOLONG and OOLONG-Pairs, motivating evaluation of effective context length as a joint property of input length and task complexity.
- The scaling experiment ranges from 2^13 to 2^20 input tokens; GPT-5’s stated context window is 272K tokens.
- Compaction repeatedly summarizes accumulated context, but can discard details needed for subsequent dense aggregation.
Strong single-needle retrieval alongside rapidly degrading dense-task scores shows why context-window size alone is an insufficient capability measure. RLM(GPT-5) preserves substantially higher OOLONG-Pairs performance and continues operating beyond the base model’s 272K-token window, although its dense-task scores still decline at longer lengths.
RLMs retain a string-input, string-output interface while moving the prompt into a Read-Eval-Print Loop (REPL) environment. Figure 2 illustrates how the root model uses symbolic access to the prompt, persistent buffers, and sub-calls instead of copying the entire input and every intermediate result into its own history.
- The authors distinguish programmatically generated sub-calls from delegation that requires the root model to verbalize each request individually.
- The implementation is available at https://github.com/alexzhang13/rlm.
The diagram locates the long prompt and intermediate values in the environment rather than the root model’s neural context. Code can store and combine sub-model responses without copying all text into the root history, though individual neural calls and the accumulating root history remain bounded.
2 Recursive Language Models
| For a base model M with maximum context size K, the framework accepts an arbitrary-length prompt P ∈ Σ⋆ and returns Y ∈ Σ⋆ through a persistent environment E. Its design goals are to process |P| ≫ K, assemble outputs beyond a single model generation, and support (\Omega( | P | )) or (\Omega( | P | ^2)) semantic work through programmatically generated calls. |
- Algorithm 1 initializes the REPL with the prompt and a sub-RLM function, then repeatedly generates and executes code until the environment contains Final.
- The root initially receives prompt metadata and subsequently receives bounded stdout observations; long strings and intermediate values remain in the environment.
- The root history still grows through generated code and observations, so the ability to launch many sub-calls does not remove practical iteration or resource limits.
The comparison with Algorithm 2 identifies three design choices: symbolic access to the input, environment-based output assembly, and sub-model calls available inside executable code. A scaffold that instead places P in model history, emits the answer directly, and exposes delegation only as a separate action retains input, output, and delegation bottlenecks.
- The concrete implementation uses a Python REPL with model-query functions available as modules and truncated printed output.
- Although the formal algorithm returns an environment variable, the experimental implementation also permits a directly generated final answer; this flexibility becomes a failure mode in some trajectories.
3 Scaling Long Context Tasks
The evaluation varies both prompt length and the information required to answer a query. A fixed-needle retrieval benchmark does not characterize tasks that require semantic processing of nearly every entry or enumeration of many qualifying pairs.
- S-NIAH requires approximately constant information, OOLONG requires linear semantic labeling and aggregation, and OOLONG-Pairs requires potentially quadratic pair output.
- The quadratic output requirement does not necessarily require quadratic neural processing: trajectories can classify entries with sub-calls and enumerate pairs in Python.
- BrowseComp-Plus and CodeQA add document-research and repository-understanding settings.
3.1 Tasks
S-NIAH contains 50 single-needle tasks that ask for a phrase or number in unrelated text. BrowseComp-Plus uses 150 randomly sampled questions, each supplied with 1000 randomly selected documents guaranteed to include the gold and evidence documents; performance is measured as the percentage of correct answers.
- BrowseComp-Plus requires multi-hop evidence assembly across several documents, although the required evidence set remains small relative to the corpus.
- The main BrowseComp-Plus inputs span 6M–11M tokens.
OOLONG uses 50 tasks from trec_coarse that require inferring semantic labels and aggregating them across nearly all questions. OOLONG-Pairs introduces 20 queries requiring lists of qualifying user-ID pairs, evaluated with F1; LongBench-v2 CodeQA asks multiple-choice questions about code repositories.
- Table 1 evaluates OOLONG at 131K tokens and OOLONG-Pairs at 32K tokens.
- CodeQA inputs range from 23K to 4.2M tokens.
- Pair enumeration tests dense input processing and long-output assembly, rather than merely counting qualifying pairs.
3.2 Methods and Baselines
The main model settings are GPT-5 with medium reasoning and default sampling parameters, and Qwen3-Coder-480B-A35B with its published sampling parameters. Baselines include direct model calls, CodeAct with BM25 or sub-calls, iterative compaction, OpenCode, and Claude Code, with coding-agent variants that either receive the context directly or access it through a file.
- The BM25 retriever is provided for BrowseComp-Plus; Appendix C.1 removes retrieval for the other tasks.
- GPT-5 compaction uses GPT-5-nano for summaries and GPT-5 for the final answer to limit the cost of ingesting millions of tokens.
- Claude Code v2.0.0 uses Claude Opus 4.1, changing both the model and scaffold. Costs use OpenAI, Fireworks, and Anthropic pricing for the corresponding models; OpenCode costs are unavailable in Table 1.
RLM depth=0 permits external context manipulation but no sub-calls; depth=1 permits calls to ordinary LLMs, and greater depths permit calls to RLMs. GPT-5 experiments use GPT-5 at the root and GPT-5-mini for recursive calls, making the scaffold a heterogeneous inference system rather than repeated use of a single model.
- The evaluated maximum recursion depths are 0–3, with depth=1 assumed when no depth is stated.
- The main text describes RLM-Qwen3-8B training on 1,000 filtered LongBenchPro trajectories generated by an RLM using Qwen3-Coder-480B-A35B with Qwen3-8B sub-calls.
- Training targets the root’s ability to manipulate the REPL and decide when to delegate; leaf requests are treated as ordinary reasoning tasks.
4 Results and Discussion
Table 1 shows that RLMs can process inputs beyond the base model’s window and improve dense processing within it. At depth=1, GPT-5 scores 62.0 on CodeQA, 91.3 on BrowseComp-Plus, 56.0 on OOLONG, and 58.0 F1 on OOLONG-Pairs, compared with direct-model scores of 24.0, 0.0, 44.0, and 0.1 respectively.
- The direct CodeQA and BrowseComp-Plus results are marked as encountering context limits; they are not clean measures of reasoning quality on fully ingested inputs.
- On BrowseComp-Plus, depth=1 RLM exceeds the compaction agent’s 70.5 and CodeAct with BM25’s 51.0, but not context-offloaded OpenCode’s 94.0.
- On OOLONG-Pairs, GPT-5 depth=1 RLM costs $0.33 ± $0.20 per task, versus $0.75 ± $0.43 for CodeAct with BM25 and $1.11 ± $0.62 for CodeAct with sub-calls.
The comparison separates context offloading from recursive delegation and reveals model-by-task interactions. GPT-5’s deeper variants improve OOLONG-Pairs, while Qwen3-Coder achieves its best CodeQA score without sub-calls; context-offloaded OpenCode exceeds the GPT-5 RLM variants on BrowseComp-Plus. Context-limit markers and unavailable OpenCode costs prevent a simple claim of overall superiority.
External context storage accounts for a substantial part of the benefit, while recursive semantic processing adds task-dependent gains. GPT-5’s OOLONG-Pairs F1 increases from 43.9 at depth=0 to 58.0, 65.5, and 76.0 at depths 1, 2, and 3, whereas Qwen3-Coder’s best CodeQA score is 66.0 at depth=0.
- Qwen3-Coder’s OOLONG score falls from 48.0 at depth=1 to 26.0 at depth=2, demonstrating that deeper recursion is not a generally monotonic improvement.
- GPT-5 CodeQA peaks at depth=2 with 66.0 and drops to 58.0 at depth=3.
- The authors attribute some Qwen3-Coder degradation to syntax errors that propagate into recursive sub-RLM calls.
Cost comparability depends on the task and baseline. GPT-5 depth=1 RLM costs $0.99 ± $1.22 on BrowseComp-Plus versus $0.57 ± $0.10 for compaction, and $0.43 ± $0.85 on OOLONG versus $0.14 ± $0.02 for the base model.
- The paper’s $1.50–$2.75 estimate for GPT-5-mini ingesting 6M–11M tokens is a linear extrapolation, not an executable full-context baseline.
- Outlier trajectories can raise average costs substantially above median costs; the distribution plots should not be read as establishing uniformly cheaper inference.
- Main-experiment LM calls are blocking and sequential. Runtime also depends on implementation, request latency, and hardware, so comparable API cost does not imply comparable latency.
Table 2 extends the evaluation to LongCoT-mini using GPT-5.2. Overall solve rate rises from 38.7 for the base model to 50.6 for depth=1 RLM and 65.6 when the RLM receives explicit decomposition hints, corresponding to the paper’s reported 69.5% relative increase for the hinted variant.
- Without hints, performance is uneven: MATH drops from 26.0 to 5.6 and CS from 40.4 to 11.0, while LOGIC reaches 86.7 and CHESS 93.0.
- With hints, scores are 32.0 for MATH, 52.0 for CHEM, 46.0 for CS, and 99.0 for both LOGIC and CHESS.
- The LongCoT-mini experiment uses a different harness, and the hinted variant adds task-specific decomposition guidance; it is distinct from the task-agnostic main benchmark configuration.
The overall gain conceals pronounced domain differences: an unhinted RLM improves LOGIC and CHESS but underperforms the base model on MATH and CS. Explicit decomposition hints improve every domain relative to the base model, showing that the orchestration guidance materially changes the result rather than leaving the scaffold configuration unchanged.
Figure 3 distinguishes gains from adding the RLM scaffold from gains after post-training. Untrained versus post-trained Qwen3-8B in the scaffold scores 26.00 versus 32.00 on CodeQA, 2.00 versus 14.00 on BrowseComp-Plus, 24.00 versus 32.04 on OOLONG, and 4.26 versus 5.17 on OOLONG-Pairs.
- Training targets orchestration, but these aggregate results do not isolate orchestration from other changes in model behavior. The post-trained 8B model remains below the frontier RLM results.
- A separate Qwen3-4B-Instruct-0527 RLVR experiment trains on MRCRv2 with 2 needles at 32K–64K tokens and evaluates on 8 needles at 512K–1M tokens.
- The MRCRv2 learning curves support length and needle-count generalization within this synthetic task, not extrapolation across arbitrary long-context problems.
The left panel distinguishes gains from scaffolding from additional gains after training the root model, with improvements on all four main tasks but low absolute scores on BrowseComp-Plus and OOLONG-Pairs. The right panel shows MRCRv2 evaluation performance increasing on the longer, higher-needle-count split during training on the shorter split, supporting generalization within this synthetic benchmark.
5 Analyses of RLM Trajectories
Observed RLM trajectories typically probe the input, choose a decomposition, and delegate selected semantic work. Sparse research tasks often use regex searches informed by model priors, while dense tasks use sub-calls for labeling and then perform aggregation or pair construction in Python.
- Persistent variables allow the system to concatenate outputs without requiring the root model to regenerate every entry.
- Some trajectories over-verify, repeat completed computations, or fail to return an already assembled answer.
Figure 4 relates outcomes to the first decomposition attempt and execution reliability. In-context examples improve OOLONG performance even when unrelated to the actual task, while Qwen3-Coder trajectories contain substantially more syntax or format errors than GPT-5 trajectories.
- A correct decomposition can still fail through implementation mistakes or erroneous sub-call results.
- An initially incorrect decomposition can sometimes be repaired through execution feedback, so first-attempt quality is influential but not determinative.
- The error comparison supports the authors’ explanation of recursion-depth failures descriptively; it is not an isolated causal intervention.
The decomposition ablation shows that examples change the initial strategy and final success, while the outcome categories distinguish planning failures from implementation and sub-call failures. The error-rate panel shows that even successful Qwen3-Coder trajectories frequently contain errors, documenting recovery while also supporting concern about error propagation at greater recursion depth.
6 Related Works
The paper places RLMs among inference scaffolds rather than architectural methods for extending context windows. It distinguishes the approach from lossy compaction, explicit memory hierarchies, human-designed delegation workflows, and one-shot programs containing sub-model calls.
- The claimed distinction from recursive delegation systems is symbolic manipulation of the user input itself, including inputs that exceed the base model’s window.
- Persistent execution feedback allows an RLM to revise code and decomposition decisions, unlike a program generated only once.
- The scaffold does not require replacing the underlying neural architecture and can also operate around longer-window models.
7 Limitations and Future Work
The authors identify limited evaluation on difficult natural long-context tasks, underexplored guardrails, and potentially exploding sub-call costs as unresolved limitations. Asynchronous execution and sandboxed REPLs are proposed implementation directions, but their benefits are not established by the sequential main experiments.
- Small models can struggle when they lack sufficient coding ability, and reasoning models can exhaust their output budget before completing REPL actions.
- Termination via FINAL() and FINAL_VAR() is brittle: a model can return a plan prematurely or fail to expose a valid stored result.
- The small-scale training evidence does not establish a scaling law for native RLM training.
8 Conclusion
The paper concludes that code-mediated access to external context and programmatic model delegation form an effective inference framework for the evaluated long-context and reasoning tasks. The results support separating persistent symbolic state from bounded neural context, while leaving training recipes, recursion policies, guardrails, and runtime implementation open.
Appendix
- A Additional Training Details describes collecting 2250 candidate trajectories from 750 English LongBenchPro tasks, retaining 1,072 after removing zero-score and single-turn trajectories, and converting root turns into supervised examples. Further filtering removes turns exceeding an approximate 100k-character limit and repairs template mistakes; 16% of turns misuse FINAL and 13% misuse FINAL_VAR. Fine-tuning uses prime-rl, batch size 64, 300 steps, and 48 H100 hours. The main text’s 1,000-trajectory description is approximate: Figure 5 reports 953 post-filtering trajectories and 4724 turn samples, whereas Appendix A reports 1,072 at an intermediate filtering stage.
- Figure 6 in A Additional Training Details reports average Qwen3-8B RLM runtimes falling from 272 to 86 seconds on CodeQA, 341 to 101 on BrowseComp-Plus, 12449 to 1649 on OOLONG, and 5113 to 532 on OOLONG-Pairs after post-training. The subsequent MRCRv2 experiment uses reinforcement learning for 150 steps, batch size 128, and 4 rollouts per example on the 32k–64k-token, 2-needle split. Each turn permits 4096 output tokens, with at most 20 RLM iterations; evaluation occurs every 50 steps starting at step 0 on the 512K–1M-token, 8-needle split.
- B Negative Results: Things We Tried That Did Not Work. documents model-specific prompting, inadequate coding ability, output-budget exhaustion, sequential-call latency, and brittle final-answer handling. Reusing the GPT-5 prompt with Qwen3-Coder caused excessive delegation, requiring a batching warning. Qwen3-235B-A22B improved from 30% to 38% on OOLONG, but some trajectories exhausted the output budget through thinking tokens before completing their actions.
- C Additional Methods and Baseline Details and C.1 Prompts for Experiments provide experimental prompt templates and model-specific changes. Interfaces include context, llm_query, rlm_query, FINAL(), and FINAL_VAR(); deeper recursion uses rlm_query(context, query), falling back to llm_query at the depth limit. Main-task RLM prompts remain fixed across tasks, but Qwen3-Coder receives a batching warning and Qwen3-8B receives adjustments for its approximately 32k-token window. CodeAct receives BM25 retrieval for BrowseComp-Plus and a prompt without retrieval for other tasks.
- C.2 Summary agent baseline specifies iterative ingestion, summarization when full, and chunking for inputs larger than the model window. GPT-5-nano performs the GPT-5 baseline’s summarization; a 20-sample BrowseComp-Plus check found comparable performance to GPT-5 summarization. This choice affects cost comparisons: the Qwen3-Coder summary agent is reported as nearly 15× more expensive on BrowseComp-Plus.
- C.3 LongCoT-mini experiment. uses Prime Intellect’s rlm-harness with sandbox execution and a different final-answer mechanism. The decomposition guidance describes planning a dependency graph, solving ready nodes through llm_batch, caching verified values, and assembling the result from stored answers; the harness scores the answer written to /task/answer.txt. Table 3 adds a base-model control in which decomposition hints reduce GPT-5.2’s overall solve rate from 38.7 to 28.6, despite improving MATH from 26.0 to 37.0.
- D Additional Benchmark Details and D.1 OOLONG-Pairs Benchmark specify 20 synthetic queries at input lengths [1024, 2048, 4096, 8192, 16384, 32768, 65536, 131072, 262144, 524288, 1048576]. Queries require inferring one of six semantic labels and returning qualifying user-ID pairs, sometimes subject to date or count constraints; explicit pair output prevents replacing enumeration with a count alone. D.2 Scaling Huge Document Corpora in BrowseComp+ evaluates 20 of the original 150 questions and adds ReAct w/ GPT-5 + BM25 and GPT-5 + pre-query BM25 baselines. At 1000 documents, RLM(GPT-5) reaches 100% and the no-recursion ablation reaches 90%; these subset results should not be conflated with the larger main evaluation.
- E Additional RLM Trajectories gives successful and failed examples. E.1 RLM(GPT-5) on BrowseComp-Plus-Query_74 finds Maria Dalmacio in an 8.3M-token corpus for $0.079 using regex probing, extraction, and verification calls. E.2 RLM(Qwen3-Coder) on OOLONG-Pairs-Query_3 costs $1.12, assembles a candidate answer whose correctness depends on its semantic classifications, repeatedly reconstructs it, then fails by generating an incorrect response instead of returning the stored result. E.3 RLM(Qwen3-Coder) on OOLONG-Query_212 reaches the correct aggregate comparison for $0.38 with excessive sub-calling, while E.4 RLM(GPT-5) on CodeQA-Query_44 answers choice 1 correctly over a 900k-token repository for $0.27.
- F Additional Quantitative Results and F.1 Additional Quantitative Analysis of Main Results compare instance-level success overlap and sub-call counts, showing that RLMs solve additional instances but do not uniformly dominate every baseline. Correct Qwen3-Coder OOLONG trajectories average about 500 sub-calls, and untrained Qwen3-8B often delegates more on incorrect trajectories. F.2 Additional Runtime and Cost Analysis of RLMs reports high-variance, long-tailed distributions, with infrequent very long executions attributed largely to sequential sub-calls. It also supplies average API costs for the input-length scaling experiment and discusses timeout logic and asynchronous calls as implementation remedies rather than demonstrated main-experiment improvements.
Brief Thoughts
The strongest contribution is the separation of neural context from persistent symbolic state, not recursion alone. The depth=0 ablation and context-offloaded coding agents show that external input access already provides substantial gains, while OOLONG-Pairs gives the clearest evidence for combining semantic sub-calls with programmatic aggregation and output construction.
The results make decomposition and execution reliability important evaluation targets, but do not establish universally cheaper or arbitrarily reliable long-context reasoning. Model-dependent depth effects, heterogeneous GPT-5/GPT-5-mini execution, unavailable OpenCode costs, and task-specific LongCoT-mini hints limit aggregate comparisons; the failure trajectories motivate explicit cost ceilings, reliable termination, and evaluation on natural dense tasks.