Rethinking the Evaluation of Harness Evolution for Agents | Summary
14 Jul 2026 | Paper Review Harness Optimization LLM Evaluation Test-Time ScalingContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Rethinking the Evaluation of Harness Evolution
- 4 Experiments
- 5 Discussion
- 6 Conclusions
- Appendix
- Brief Thoughts
This article explains the key points of Rethinking the Evaluation of Harness Evolution for Agents.
- 2026-07-14 (arXiv)
- Wang, Yike, Zhu, Huaisheng, Hu, Zhengyu, Yuan, Yige, Chen, Zhengyu, Senthil, Shakti, Hajishirzi, Hannaneh, Tsvetkov, Yulia, Dasigi, Pradeep, Xiao, Teng.
- Allen Institute for AI, University of Washington, Independent
- Paper
- Github
- Project Page
Summary
- Automatic harness evolution revises an LLM agent’s external scaffold using task feedback. When search and final evaluation use the same benchmark, a higher score may reflect repeated search over those tasks rather than a reusable harness improvement.
- The authors compare parallel sampling, sequential refinement, shared harness evolution, and task-specific harness scaling under a common rollout budget, both with and without unit-test feedback. They separately evolve a harness on training tasks, select it on validation tasks, and evaluate it on held-out Terminal-Bench 2.1 tasks.
- Harness evolution does not consistently outperform the simpler search baselines. On held-out tasks evaluated with Claude Opus 4.6 and GPT-5.4, its average pass@1 rises only from 67.7 to 68.3, motivating separate search and evaluation tasks.
1 Introduction
A harness specifies how a fixed model receives instructions, uses tools, retains information, and controls execution. Because harness-search methods inspect trajectories and feedback from benchmark tasks, the paper asks whether their reported gains exceed those from spending additional attempts directly on the tasks, and whether the resulting harness transfers.
The dashed initial-harness reference is 68.2. Without unit-test feedback, Parallel Sampling leads at 72.3, followed by Harness Scaling at 71.8 and Sequential Refinement at 69.3; shared Harness Evolution falls below the reference at 67.4.
2 Related Work
Test-time scaling allocates more inference computation without changing model weights: parallel sampling explores independent candidates, while sequential refinement builds on previous attempts. Earlier optimizers target prompts or stored experience; Meta-Harness, Agentic Harness Engineering (AHE), and AEVO edit broader agent scaffolds. This paper evaluates harness evolution rather than introducing another general-purpose optimizer.
3 Rethinking the Evaluation of Harness Evolution
The four methods allocate additional attempts to independent trajectories, successive revisions, a harness shared across tasks, or a harness revised for one task. The framework distinguishes a frozen harness from an evolving one and identifies the trajectories, summarized experience, and feedback available to the meta agent that proposes edits.
The diagram distinguishes independent samples and trajectory refinement under a frozen harness from a harness evolved across a task batch and one revised separately for each task. The latter two use a meta agent, but only shared evolution aims to produce a reusable harness.
3.1 Preliminary
For task x, policy πθ generates a trajectory y under harness h; when a unit test g is available, R(y, g) indicates success. Each method has budget K and returns a final trajectory. A summarization map Φ, realized by the underlying model, makes stored attempts and outcomes available to subsequent refinement or harness editing.
3.2 Parallel Sampling
Parallel Sampling draws K independent trajectories with the same harness h. Without unit tests, an agent self judge selects a candidate; with tests, the method can return any accepted trajectory. This baseline measures the benefit of additional attempts without harness revision.
3.3 Sequential Refinement
Sequential Refinement keeps h fixed but conditions each later trajectory on a summary of the previous one. When unit tests are available, that summary also includes the previous outcome. The method returns its final attempt without tests or an accepted attempt when tests permit selection.
3.4 Harness Evolution
Harness Evolution runs tasks under a shared harness, stores their trajectories and, when available, unit-test outcomes, then asks a meta agent to revise that harness using a summary of accumulated experience. It uses the last harness without unit tests or selects one by measured aggregate outcome when tests are available before running the target task. The AHE implementation disables its explore agent to prevent retrieval of externally available benchmark-specific harnesses.
3.5 Harness Scaling
Harness Scaling revises the harness separately for each evaluation task. Its meta agent conditions on that task, the previous harness and trajectory, and any available unit-test outcome. Unlike shared Harness Evolution, this procedure spends revisions on instance-specific adaptation rather than producing a harness intended for reuse.
4 Experiments
The experiments compare methods without unit-test feedback, repeat the comparison with unit-test feedback, and test a selected evolved harness on tasks excluded from search. The first two settings assess uses of additional attempts; the third measures transfer.
4.1 Experimental Setup
Terminal-Bench 2.1 retains 89 terminal tasks and repairs 28 tasks from version 2.0. The reported experiments use Claude Opus 4.6, GPT-5.4, and, where specified, GPT-5.4 mini; average two independent runs; use the same initial harness; set K = 5; and use one rollout per task per harness in AHE (m = 1). Unless noted otherwise, models use high reasoning effort and a maximum generation budget of 128k tokens.
Parallel Sampling improves on direct sampling for each of the three models. Harness Evolution averages 67.4 against the initial harness’s 68.2, including a GPT-5.4 decline from 75.3 to 69.7; task-specific Harness Scaling averages 71.8.
4.2 Without Unit Test Cases, Automatic Harness Evolution Underperforms Test-Time Scaling
Without unit-test feedback, average pass@1 across the three models is 68.2 for direct sampling with the initial harness, 72.3 for Parallel Sampling, 69.3 for Sequential Refinement, 67.4 for Harness Evolution, and 71.8 for Harness Scaling. Harness Evolution lowers GPT-5.4 performance from 75.3 to 69.7, while Harness Scaling achieves the highest Claude Opus 4.6 score, 76.0. Parallel Sampling is the only method in this comparison that improves on direct sampling for all three models.
With unit-test access, Parallel Sampling has the highest mean pass@1, 86.0, while Sequential Refinement has the highest mean pass@5, 91.8. Harness Evolution scores 75.8 and 86.2, respectively, distinguishing average rollout success from finding at least one success within five attempts.
4.3 With Unit Test Cases, Automatic Harness Evolution Underperforms Test-Time Scaling
With unit tests available for feedback and selection, average pass@1 across Claude Opus 4.6 and GPT-5.4 is 86.0 for Parallel Sampling, 84.3 for Sequential Refinement, 75.8 for Harness Evolution, and 82.6 for Harness Scaling, versus 72.9 for direct sampling. Sequential Refinement has the highest average pass@5, 91.8, compared with 86.2 for Harness Evolution. The gap between Harness Evolution’s pass@1 and pass@5 indicates that success within repeated attempts does not establish an equally large improvement in average rollout success.
After evolution on training tasks and selection on validation tasks, held-out pass@1 improves by 1.2 points for Claude Opus 4.6 and 0.0 for GPT-5.4. The average is 68.3 versus 67.7 for the initial harness, measuring transfer to disjoint tasks within Terminal-Bench 2.1.
4.4 Harness Evolution Fails to Generalize Beyond Training Tasks
The authors divide the 89 tasks into 45 training, 10 validation, and 34 held-out test tasks, evolve with unit tests on training tasks, and select a harness using validation performance. Held-out pass@1 changes from 63.3 to 64.5 for Claude Opus 4.6 and remains 72.1 for GPT-5.4; the two-model average changes from 67.7 to 68.3 (+0.6). This measures limited transfer within the benchmark, without identifying which individual edits do or do not transfer.
5 Discussion
The authors inspect edit trajectories to explain why plausible prompt, middleware, and tool changes yield limited aggregate gains. They distinguish edits that save time or prevent repeated setup failures from changes that would enable solutions to persistently hard tasks.
5.1 Rational Harness Edits but Marginal Gains
Shared Harness Evolution adds prompt and memory guidance, then may introduce turn-budget reminders, tool-output truncation, finalization gates, and tool guidance. Harness Scaling more often stores task-specific bugs, paths, command plans, and checks observed in earlier attempts. The authors report that such edits can avert avoidable failures but frequently preserve facts specific to a task, while added prompt text can consume context.
5.2 Task Difficulty and Harness Sensitivity as Factors
The discussion proposes two explanations for weak gains on Terminal-Bench: relatively high baseline scores may leave failures rooted in model capability, and many solvable tasks may need little beyond a basic prompt and shell tool. Neither explanation is experimentally isolated. The authors suggest testing tasks with both substantial remaining headroom and a consequential dependence on harness components.
6 Conclusions
The conclusion concerns the evidence for harness evolution, not whether harness design can ever help. Reusing benchmark tasks for search and final measurement risks conflating reusable improvement with task-specific discovery; under the paper’s comparisons, simple test-time baselines remain competitive and held-out gains are small.
Appendix
- A Experimental Details: A.1 Initial Harness specifies an identical minimal code-agent harness for all four methods, with one bash tool and no skills, middleware, or persistent memory. In A.2 Summarization Map, AHE’s Agent Debugger organizes traces as files, writes task-level analyses, and aggregates them into an overview while keeping original traces accessible. A.3 Agentic Harness Engineering (AHE) disables the explore agent, retaining the propose-and-refine loop without externally retrieved benchmark-specific harnesses.
- A.4 Configuration states that each rollout uses a fresh E2B remote sandbox. Table 4 specifies high reasoning effort, a 200,000-token maximum context, and a 128,000-token maximum generation per turn for the code agent, debugger, and meta agent, with maximum model turns of 300, 25, and 500, respectively. A.5 Evaluation Metrics defines pass@1 as mean binary reward over rollouts and pass@k as the fraction of tasks with at least one successful rollout; infrastructure exceptions count as failures.
- B Case Study gives task-level Harness Scaling edits. Figure 3 includes compile-compcert passing after guidance on Coq 8.16.1, shell timeouts, and make install; db-wal-recovery passing after instructions to back up database and WAL files before opening SQLite; and mteb-retrieve passing at iteration 5 after distinguishing query and passage embedding prompt types. These examples show both concrete recoveries and the task specificity of some stored guidance.
Brief Thoughts
The fixed-harness search baselines and held-out split address distinct evaluation questions: whether extra attempts explain gains, and whether an evolved harness transfers. The transfer result covers one benchmark and two models, with results averaged over two runs. Matching K and AHE’s rollout count does not establish identical total token expenditure when methods also invoke a debugger or meta agent.