Meta-Harness: End-to-End Optimization of Model Harnesses | Summary
30 Mar 2026 | Paper Review Harness Optimization AI Coding AgentsContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Meta-Harness: A Harness for Optimizing Harnesses
- 4 Experiments
- 5 Discussion
- Appendix
- Brief Thoughts
This article explains the key points of Meta-Harness: End-to-End Optimization of Model Harnesses.
- 2026-03-30 (arXiv)
- Lee, Yoonho, Nair, Roshen, Zhang, Qizheng, Lee, Kangwook, Khattab, Omar, Finn, Chelsea.
- Stanford, KRAFTON, MIT
- Paper
Summary
- Meta-Harness: End-to-End Optimization of Model Harnesses searches over the executable code surrounding a frozen language model. Its coding-agent proposer selectively reads a filesystem containing all previous candidates’ source code, evaluation scores, and execution traces, rather than receiving only scalar rewards or predetermined summaries.
- On online text classification, the selected harness achieves 48.6% test accuracy versus ACE’s 40.9%, with 11.4K versus 50.8K additional context tokens. A discovered math retrieval harness improves average accuracy from 34.1% to 38.8% across five evaluated models on 200 IMO-level problems; TerminalBench-2 harnesses achieve 76.4% with Claude Opus 4.6 and 37.6% with Claude Haiku 4.5.
- The classification interface ablation provides the clearest evidence for raw trace access: median candidate search accuracy is 50.0% with the full interface, compared with 34.6% for scores and code alone and 34.9% when summaries are added. Unseen classification datasets and four math models not used during search provide transfer evidence, but TerminalBench-2 uses the same tasks for search and final evaluation, and alternative proposer agents remain untested.
The left panel summarizes evaluation-count efficiency in classification, while the right compares Haiku 4.5 harness pass rates. The classification curve measures search-set progress; the TerminalBench-2 bars report benchmark performance after search on that same benchmark.
1 Introduction
A model harness determines what information is stored, retrieved, and presented to an LLM, and how interactions and state updates are orchestrated. The paper asks whether the manual process of inspecting failures and revising this code can be automated without changing model weights.
- Long-horizon failures may originate in earlier retrieval, memory, or prompt-construction decisions; aggregate scores alone do not expose these dependencies.
- Table 1 compares estimated diagnostic context generated per artifact evaluation, not tokens necessarily consumed by the proposer. The largest Meta-Harness setting generates up to 10,000,000 tokens, motivating selective filesystem access.
- Project page: https://yoonholee.com/meta-harness/. Optimized TerminalBench-2 artifact: https://github.com/stanford-iris-lab/meta-harness-tbench2-artifact.
Estimated context generated per evaluation ranges from 0.002–0.026 million tokens for the listed prior settings to 10.0 million for Meta-Harness. The table describes available diagnostic scale, not actual proposer consumption or computational efficiency under a common task.
2 Related Work
Meta-Harness connects harness-level credit assignment with adaptive external-context access, executable program search, and text optimization. It uses past rollouts to diagnose failures and rewrite the external procedure governing subsequent behavior, rather than updating model weights.
External memory and adaptive access.
Retrieval-augmented generation, interleaved retrieval and reasoning, memory-based agents, and recursive language models treat large information sources as resources to query adaptively. Meta-Harness applies this access pattern to an optimization history containing code, scores, and execution traces, allowing the proposer to revise the context-management procedures themselves.
Executable code search.
Prior executable-code search methods evolve functions, agents, workflow graphs, or memory designs. Meta-Harness allows edits to complete domain-specific harness implementations, including prompting, retrieval, and state updates, without imposing a fixed parent-selection rule or predefined mutation operators.
- The outer-loop search history persists across candidates, while task-specific harness state resets between tasks in the setting distinguished by the paper.
- The proposer chooses which prior artifacts and failures to inspect through access to the experience filesystem; the search procedure does not prescribe a single parent’s feedback.
Text optimization methods.
ProTeGi, TextGrad, OPRO, GEPA, AlphaEvolve/OpenEvolve, and Feedback Descent optimize text or code using feedback from prior attempts. The paper argues that their usual feedback interfaces compress or structure information before its diagnostic relevance is known, whereas harness engineering may require comparing raw traces across examples and candidates.
- Table 1 summarizes the settings studied in the cited papers; it is not a controlled performance comparison.
- The classification experiments directly compare Meta-Harness with GEPA, Best-of-N, OpenEvolve, and the text-optimization component of TTT-Discover.
3 Meta-Harness: A Harness for Optimizing Harnesses
Meta-Harness is an outer-loop harness that repeatedly proposes, evaluates, and logs task-specific harnesses. It retains prior experience in a queryable filesystem and delegates diagnosis and program revision to a coding agent.
Objective.
For a frozen model M, a task distribution X, and a stateful harness H, a rollout follows the harness’s prompt-construction and state-update rules. The optimization objective is H* = arg max_H E_{x∼X, τ∼p_M(H,x)} r(τ,x), where r scores the resulting trajectory; when accuracy and context cost both matter, candidates are compared through Pareto dominance.
Meta-Harness search loop.
Each evaluated candidate contributes a directory containing its source code, scores, and execution traces, including prompts, tool calls, model outputs, and state updates. The proposer queries this growing history with tools such as grep and cat, diagnoses likely failure modes, and produces one or more new harnesses.
- Algorithm 1 evaluates the initial population, then validates each proposed harness against the interface before running an expensive evaluation and adding its artifacts to the filesystem.
- The search retains evaluated candidates and returns a Pareto frontier, but the proposer can inspect any prior candidate rather than being assigned a parent.
- Classification and math candidates are selected using search-set feedback without exposing test-set results to the proposer. TerminalBench-2 is an explicitly disclosed exception because search and final evaluation use the same benchmark.
The feedback path retains proposed code, reasoning traces, and scores as filesystem artifacts. The proposer can revisit them selectively, making retrieval of prior experience part of optimization rather than a fixed prompt-construction step.
Advantages of code-space search.
Code-space search permits changes to algorithmic structure, from retrieval and memory rules to full rewrites, rather than only template filling or local prompt mutations. The qualitative traces show the proposer comparing failed edits, identifying a shared prompt intervention, and eventually choosing an additive modification after repeated control-flow regressions.
- The proposed regularization benefit—coding models tend to generate coherent, reusable algorithms—is an interpretation rather than a separately measured effect.
- The TerminalBench-2 narrative documents hypothesis-driven diagnosis, but does not experimentally establish that each causal explanation offered by the proposer is correct.
Practical implementation.
Experimental harnesses are single-file Python programs modifying domain-specific prompting, retrieval, memory, and orchestration. The proposer is Claude Code with Opus-4.6, guided by a minimal skill specifying artifact locations, permitted edits, and how to inspect prior runs; the underlying task model remains frozen.
- The paper describes a typical run as roughly 60 evaluated harnesses over 20 iterations, while reporting domain-specific budgets separately.
- Interface validation screens proposed programs before full benchmark evaluation, and candidate scoring occurs outside the proposer.
4 Experiments
The experiments cover online text classification, retrieval-augmented mathematical reasoning, and autonomous terminal tasks. Comparisons include hand-designed domain harnesses and program-search methods, with accuracy, context cost, or pass rate reported according to the setting.
4.1 Online Text Classification
Online classification uses GPT-OSS-120B: the model receives labeled examples sequentially, updates memory, and predicts held-out examples. Search covers LawBench with 215 classes, Symptom2Disease with 22 classes, and USPTO-50k with 180 classes, starting from zero-shot, few-shot, ACE, and MCE harnesses and producing 40 candidates over 20 iterations with two candidates per iteration.
Comparison vs text optimizers. The paper uses the same Opus-4.6 proposer configuration with max reasoning, an equal budget of candidate harness evaluations, and search-set-only selection. Table 4 reports median/best candidate search accuracy of 32.6/40.2 for GEPA, 34.0/44.2 for Best-of-N, 39.1/43.3 for OpenEvolve, 34.1/45.6 for TTT-Discover, and 50.0/56.7 for Meta-Harness.
- Best-of-N independently samples from the seed; the TTT-Discover comparison uses its text-optimization component with the PUCT reuse rule, not the complete method.
- Figures 1 and 4 report that Meta-Harness reaches OpenEvolve’s and TTT-Discover’s final accuracy within four evaluations and finishes more than 10 points ahead.
- “Meta-Harness is 10× Faster and Converges to a Better Harness” refers to full evaluation count, not a demonstrated 10× reduction in wall-clock time or total proposer-token cost.
Meta-Harness has the highest median and best candidate search accuracy. Its 50.0% median exceeds every comparator’s best score, so the reported advantage is not confined to one exceptional candidate; the table does not quantify variation across independent searches.
Candidate dots and best-so-far lines distinguish proposal quality from the best retained result. Meta-Harness reaches 56.7%, whereas the strongest comparator finishes at 45.6%; these are search-set results, not held-out test accuracies.
The interface ablation in Table 3 retains code and scores in all conditions. Scores Only achieves 34.6% median and 41.3% best candidate accuracy; Scores + Summary achieves 34.9% median and 38.7% best; the full interface with raw execution traces achieves 50.0% median and 56.7% best.
- The numbers of evaluated candidates exceeding the zero-shot baseline are 26, 23, and 39, respectively.
- The full condition’s median exceeds either restricted condition’s best result, supporting the value of raw trace access in this classification setting.
- The summary condition does not recover the lost diagnostic information; its lower best score alone does not establish that summaries are generally harmful. These medians describe candidates, not variability across independent search repetitions.
Code and scores remain available in every condition, while the full interface adds raw trace access. Its median candidate accuracy exceeds the best result under either restricted interface; generated summaries do not close the gap.
Comparison vs state-of-the-art harnesses. The selected harness reaches 48.6% average test accuracy, compared with 40.9% for ACE and 40.0% for MCE, while using 11.4K additional context tokens versus 50.8K and 28.5K. The gain is uneven: Meta-Harness scores 14.0% on USPTO, 86.8% on Symptom2Disease, and 45.0% on LawBench, falling below ACE on USPTO.
- Accuracy–Context Tradeoffs. Figure 3 shows multiple discovered operating points rather than a single fixed context budget. Its horizontal axis measures additional context in characters, although the caption describes tokens; Table 2 separately reports tokens.
- Table 9 reports eight frontier variants in thousands of additional characters, ranging from Draft Verification at 40.1% accuracy and 5.4K characters to Label-Primed Query at 48.6% and 45.5K characters.
- Figure 7 compares search and test accuracy per dataset. The spread around the diagonal, especially for USPTO, shows that search-set gains do not translate uniformly into test gains.
The selected harness improves average accuracy through Symptom2Disease and LawBench, but falls below ACE on USPTO. Its 11.4K additional tokens are substantially below ACE’s 50.8K, so the average improvement does not require a larger input context.
The discovered operating points offer higher accuracy across a range of context budgets, rather than one fixed design. The horizontal axis explicitly uses characters, which must be distinguished from Table 2’s token units and the caption’s token wording.
The eight variants span different accuracy–context operating points, with Draft Verification at the lowest context cost and Label-Primed Query at the highest average accuracy. Context is measured in thousands of characters, not the thousands of tokens reported in Table 2.
The per-dataset plots show that search and test performance are related but not interchangeable. Many USPTO candidates fall below the diagonal, cautioning against treating aggregate search progress as uniform generalization.
Out-of-distribution (OOD) task evaluation. On nine datasets unseen during search, the selected harness achieves 73.1% average accuracy, compared with ACE’s 70.2% and the strongest reported few-shot average of 69.6%. It attains the highest performance on 6/9 datasets, including a tie, and uses 7.3K average additional context tokens versus ACE’s 11.7K.
- Table 5 includes citation intent, financial entities, review ratings, financial sentiment, emotions, banking intent, news topics, science entailment, and tweet hate speech.
- ACE is stronger on FiNER and Financial PhraseBank, and Few-shot (32) is stronger on Banking77.
- Using all available few-shot examples rather than 32 reduces accuracy on 7/9 tasks; more context is not consistently better in this evaluation.
The 73.1% average supports transfer to unseen datasets, with lower average context cost than ACE. Individual columns also show limits: ACE leads on FiNER and Financial PhraseBank, while Few-shot (32) leads on Banking77.
The discovered classification programs implement two ways to interrogate labeled memory. Draft Verification retrieves five neighbors for an initial prediction, then five same-label confirmers and five different-label challengers; Label-Primed Query constructs one prompt containing valid labels, one query-relevant example per known label, and local contrastive pairs selected using TF-IDF.
- Figure 5 shows prediction-conditioned second-stage retrieval; fewer than five stored examples trigger a single-call few-shot fallback.
- Figure 6 shows the higher-accuracy single-call policy: class coverage exposes the label space, while similar examples with different labels make local decision boundaries explicit.
- These procedures are representative endpoints of the discovered frontier, not manually specified search templates.
The second retrieval is conditioned on the draft label, allowing verification to target the current decision with confirmers and challengers. This differs from simply retrieving more generic nearest neighbors.
The single-call prompt combines global label-space coverage with locally contrasting examples. It exposes both possible answers and query-specific decision boundaries without a separate verification call.
4.2 Harnesses for Retrieval-Augmented Reasoning
Math search optimizes retrieval behavior over a corpus of 535,356 solved problems from eight datasets. The main experiment uses 40 iterations, 109 candidate harnesses, and a 250-problem search set from OlympiadBench and Omni-MATH hard, selecting one harness using GPT-OSS-20B; initial harnesses are zero-shot, few-shot, and ACE.
- Final evaluation contains 200 unseen IMO-level problems: 100 from IMO-AnswerBench, 60 from IMO-ProofBench, 17 from ArXivMath December 2025, and 23 from ArXivMath January 2026, as detailed in Table 11.
- Evaluation uses GPT-OSS-20B plus four models not seen during search: GPT-5.4-nano, GPT-5.4-mini, Gemini-3.1-Flash-Lite, and Gemini-3-Flash.
- Table 10 describes the corpus. Exact-prefix and fuzzy Jaccard filtering with threshold 0.8 remove matches against evaluation and search problems; the authors also manually inspect top BM25 retrievals.
The corpus combines 535,356 problems with heterogeneous solution lengths and proof content; its reported total proof proportion is 22%. The source statistics describe the retrieval resource, while decontamination and runtime solution-length filters further constrain usable entries.
Half of the evaluation set is the sampled IMO-AnswerBench subset; the rest combines proof and research-style benchmarks. The table provides composition, not separate performance estimates, and therefore does not identify which benchmark contributes most to the aggregate gain.
Results. Table 6 reports pass@1 averaged over three samples per problem: Meta-Harness reaches 38.8% across the five evaluated models, compared with 34.1% without retrieval, 34.4% for dense retrieval with k=1, 38.1% with k=5, 32.2% for random few-shot prompting, and 37.5% for BM25 retrieval. Its gains over no retrieval are +8.7, +1.6, +6.3, +3.7, and +3.0 points in the table’s model order.
- The discovered harness uses the same BM25-based lexical stack as the sparse baseline, whereas dense retrieval uses text-embedding-3-small.
- The average gain over no retrieval is 4.7 points, with improvements on all five evaluated models. Its average advantage over BM25 is 1.3 points.
- The highest average does not imply per-model dominance: dense retrieval with k=5 is stronger on Gemini-3.1-Flash-Lite and Gemini-3-Flash, and fixed BM25 is slightly stronger on Gemini-3-Flash.
- The statement “Meta-Harness Improves Reasoning on IMO-Level Math Problems” is supported by these results, but the PDF’s “five held-out models” wording is broader than its protocol: GPT-OSS-20B is used during search, while four other models are unseen.
The discovered math harness is a four-route BM25 program, illustrated in Figure 8. Lightweight keyword and regex predicates choose exactly one route—combinatorics, geometry, number theory, or a default algebra/other route—and only that route supplies examples to the final prompt.
- Combinatorics retrieves 20 candidates, deduplicates to eight, reranks by lexical score and difficulty, and keeps three; geometry combines one hard NuminaMath reference with two raw BM25 neighbors.
- Number theory retrieves 12 candidates and keeps three after reranking with lexical score, difficulty, and a bonus for early technique statements. The default route retrieves 10 candidates and adapts the number of examples to score concentration.
- The selected program merges two successful search lineages. Retrieval is restricted to non-empty solutions shorter than 4,000 characters, with inserted solutions truncated to 3,000 characters.
Routing is a discrete choice rather than an ensemble: exactly one subject-specific policy supplies the final context. Deduplication, reranking, and example counts differ by route and form part of the executable program selected through search.
4.3 Evaluating Agentic Coding Harnesses on TerminalBench-2
TerminalBench-2 contains 89 difficult autonomous terminal tasks. Search begins from Terminus 2 and Terminus-KIRA, but uses the same 89 tasks for search and final evaluation; the authors frame this as benchmark discovery rather than a held-out generalization experiment.
- Manual inspection and regex audits check evolved harnesses for task-specific string leakage.
- These audits address visible hard-coding, not every form of benchmark-specific adaptation or selection bias.
- The appendix analyzes one 10-iteration Claude Opus 4.6 search run; its local baseline scores should not be conflated with the final leaderboard comparison.
Results. Table 7 reports a 76.4% pass rate for Meta-Harness with Claude Opus 4.6, versus Terminus-KIRA’s 74.7%, ranking second behind ForgeCode’s reported 81.8%. With Claude Haiku 4.5, Meta-Harness reaches 37.6%, exceeding Goose’s 35.5% by 2.1 points and ranking first among the reported Haiku 4.5 agents.
- Comparator results come from the official leaderboard rather than uniformly rerun evaluations. The callout “Meta-Harness Surpasses Hand-Engineered Agents on TerminalBench-2” refers specifically to the reported model-specific comparisons.
- The authors report being unable to reproduce ForgeCode’s score from its public code alone; this does not invalidate its listed leaderboard result.
- The paper does not report confidence intervals for these pass-rate differences.
Meta-Harness ranks second among listed Opus 4.6 agents and first among listed Haiku 4.5 agents. Comparator values come from the official leaderboard, and the discovered harness was optimized on the same benchmark used for final scoring.
Qualitative behavior of the proposer. After repeated regressions from prompt and completion-flow edits, the proposer selects environment bootstrapping: a compound shell command collects a sandbox snapshot before the first model call and appends it to the initial prompt. Figure 9 separates this discovered component from inherited Terminus-KIRA features: native tool calling, a 30KB output cap, and a multi-perspective completion checklist.
- The snapshot includes the working directory, up to 20 /app entries, language versions, package managers, and available memory; a 15-second timeout and silent failure protect unusual environments.
- The implementation adds roughly 80 lines. The paper reports gains on 7 of 89 tasks, with the largest on protein-assembly and path-tracing, and associates these gains with tasks requiring non-obvious tooling.
- Table 8 records a median of 82 file reads per iteration, with 41% accessing harness source and 40% execution traces. File counts document historical artifact use but do not measure the causal contribution of each read.
The near-equal proportions of source-code and execution-trace reads document use of both implementations and observed behavior. The statistics establish broad artifact access, but do not identify which reads caused an improvement.
The discovered environment snapshot is inserted before the inherited agent loop, preserving the existing tool-use and completion machinery. This additive design follows the proposer’s shift away from prompt and control-flow changes that repeatedly regressed.
5 Discussion
The discussion emphasizes selective access to diagnostic experience rather than code search alone. OOD classification performance and transfer to unseen math models support reusable harness strategies, while readable code makes explicit hard-coding easier to inspect.
- The authors describe search runs completing in a few hours, but provide no comprehensive accounting of evaluation cost, proposer-token use, or variability across runs.
- All three domains use Claude Code with Opus-4.6 as the proposer; whether the interface advantage holds for weaker or different proposers remains open.
- Co-evolving harnesses and model weights is proposed as future work, not evaluated in this paper.
Appendix
- A Qualitative Proposer Behavior; A.1 File Access Statistics. In the analyzed 10-iteration TerminalBench-2 run, the proposer reads a median of 82 files per iteration, ranging from 69–99. Reads comprise harness code (41%), execution traces (40%), score/summary files (6%), and other files (13%), documenting access beyond the most recent candidate.
- A.2 Qualitative Behavior: Causal Reasoning Over Prior Failures. Iterations 1–2 combine structural fixes with cleanup-oriented prompt edits and regress from the run’s 64.4% baseline to 58.9% and 57.8%; iteration 3 identifies the shared prompt intervention and tests isolated fixes, reaching 63.3%. Iterations 4–6 still regress when changing completion logic, prompt language, or waiting behavior; iteration 7 pivots to additive environment bootstrapping and becomes the run’s best candidate, iteration 8 attempts composition with earlier fixes, and iteration 10 references lessons from another run after an iteration-9 candidate crashed before evaluation. The appendix’s Summary interprets this sequence as diagnosis and hypothesis revision; it is a qualitative account, not a controlled causal study.
- B Discovered Harnesses; B.1 Text Classification Harness. The appendix presents method-style abstractions of executable programs typically containing 100–1000 lines, rather than complete code listings. Its Overview contrasts Meta-Harness (Draft Verification), which uses two short calls with prediction-conditioned supporting and challenging retrieval, with Meta-Harness (Label-Primed Query), the main selected system, which uses one larger prompt combining label priming, class coverage, and query-anchored contrastive pairs. Table 9 lists the broader frontier, including Error-Annotated, CoT Replay, Cluster Coverage, Cascade Retrieval, RRF + Contrastive, and Relevance + Contrastive variants.
- B.2 Math Retrieval Harness; Overview. The lexical router selects one subject-specific BM25 policy, with route-specific candidate counts, deduplication, difficulty reranking, and example budgets. The math-aware tokenizer preserves LaTeX tokens, and the geometry hard-reference index uses NuminaMath problems with difficulty greater than 6. The final program autonomously combines geometry and combinatorics strategies from different search lineages.
- B.3 TerminalBench-2 Harness; Per-task analysis. Environment bootstrapping adds a sandbox snapshot without modifying the inherited agent loop or completion checklist. The appendix attributes its value to avoiding the first 2–4 exploratory turns on tasks with uncertain tooling, and reports gains on seven tasks, particularly protein-assembly and path-tracing. This is a task-level interpretation of observed changes, not an independently measured universal turn-saving guarantee.
- C Dataset Details; C.1 OOD Text Classification Datasets. The nine datasets are SciCite, FiNER-139, Amazon Reviews, Financial PhraseBank, GoEmotions, Banking77, AG News, SciTail, and TweetEval (Hate). They span 3-way citation intent, 139 financial numeric entity types, 5-way ratings, 3-way financial sentiment, 28 emotion labels including neutral, 77 banking intents, 4-way news topics, science entailment, and binary hate-speech detection.
- C.2 Math Retrieval Corpus. Table 10 lists OpenMathReasoning (281,743 problems), DeepMath-103K (103,021), NuminaMath-1.5 (129,520), PolyMath (11,083), Omni-MATH (4,289), FineProofs-SFT (4,275), AIME 1983–2024 (933), and Putnam-AXIOM (492), totaling 535,356. Filtering includes competition-source selection for NuminaMath, one-solution-per-problem deduplication and evaluation-family removal for OpenMathReasoning, corpus-wide exact-prefix and fuzzy Jaccard decontamination at threshold 0.8, and 5,000-character truncation for OpenMathReasoning and DeepMath solutions.
- C.3 Math IMO-level Test Set. Table 11 specifies a stratified 100-problem IMO-AnswerBench subset, all 60 IMO-ProofBench problems, and 17 and 23 problems from ArXivMath December 2025 and January 2026. The aggregate mixes answer-style, proof, and research-style problems, but the PDF provides only the count breakdown here, not per-benchmark Base-versus-Meta-Harness performance.
- D Practical Implementation Tips. The authors recommend refining the domain skill through short 3–5-iteration runs, choosing difficult and discriminative search examples, logging machine-readable artifacts, optionally providing a navigation CLI, validating interfaces before costly evaluation, and scoring candidates outside the proposer. These are engineering lessons rather than controlled findings. The advice cites 88 math problems as a practical search-set example, whereas Section 4.2 reports 250; the PDF does not reconcile these counts.
- E Extended Related Work. AlphaEvolve / OpenEvolve are contrasted with Meta-Harness through structured program databases, scalar scores, and parent/mutation mechanisms; GEPA is contrasted through candidate-level reflective feedback rather than adaptive inspection of the full population. Prompt orchestration frameworks—LMQL, LangChain, and DSPy—help specify and organize LLM programs, whereas Meta-Harness optimizes the executable implementation of retrieval, memory, and orchestration policies. These distinctions describe the paper’s positioning, not proof that the other systems cannot be adapted to richer histories.
Brief Thoughts
The clearest contribution is the feedback interface: the classification ablation supports preserving raw traces and letting the proposer decide what to inspect. The discovered policies are concrete and interpretable, but evaluation strength varies by domain: unseen datasets and models provide transfer evidence, while TerminalBench-2 demonstrates benchmark-specific improvement rather than held-out generalization. A useful follow-up would compare proposer agents, repeat searches with uncertainty estimates, and separate evaluation-count efficiency from total computational cost. The math gain is positive against no retrieval on every evaluated model but only 0.7 points above dense retrieval with k=5 on average; the coding leaderboard gains lack reported confidence intervals. The evidence supports selective access to complete diagnostic history for executable harness search in the tested settings, not the universal superiority of an unstructured outer loop.