HarnessOpt-Bench: Evaluating LLMs at Harness Optimization | Summary
06 Aug 2026 | Paper Review Harness Optimization LLM Evaluation Agent LearningContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 HarnessOpt-Bench: Harness Optimization as a Task
- 4 Experimental Setup
- 5 Results
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of HarnessOpt-Bench: Evaluating LLMs at Harness Optimization.
- 2026-08-06 (arXiv)
- Ursekar, Varun, Shanker, Apaar, Maurya, Yash, Yasser, Shehab, Kalmath, Vijay S., Chatrath, Veronica, Xue, Yuan.
- Scale AI
- Paper
Summary
- HARNESSOPT-BENCH evaluates whether an LLM operating through a coding harness can improve a target agent’s prompts, tools, control flow, memory, and orchestration code under expensive stochastic evaluation. Across 111 scored runs on four downstream tasks, optimizer-model differences are larger on average than coding-harness differences, while native coding harnesses show no consistent advantage.
1 Introduction
The paper treats harness optimization as harder than ordinary code editing because the optimizer must infer improvement from noisy and costly agent evaluations under a fixed target-evaluation budget. It holds the target model, environment, and verifier fixed; reserves final scoring for an inaccessible test partition; and uses a trusted boundary to enforce access restrictions, resource accounting, and candidate auditability.
2 Related Work
Related work includes prompt and textual-artifact optimization, evolutionary program search, and code-defined agent-design search. The closest methods use coding agents as end-to-end harness optimizers, but their reported results often vary jointly with seeds, budgets, search spaces, and evaluation protocols; HARNESSOPT-BENCH fixes these factors to support controlled comparison.
3 HarnessOpt-Bench: Harness Optimization as a Task
The benchmark defines harness optimization as constrained stochastic program optimization. An optimizer edits a pinned executable seed harness, evaluates candidate versions through graded feedback within a resource budget, and nominates one candidate for hidden test evaluation.
3.1 The optimization problem
A task fixes invariants θ = (M, E, V): the target models available to candidates, each case’s environment, and a verifier that maps trajectories to scores in [0, 1]. The optimizer may modify the feasible harness H but not θ; development feedback reveals inputs, per-case outcomes, and traces, validation reveals aggregate scores, and test feedback remains unavailable until nomination. The benchmark reports normalized gain g = (Eθ(H+) − Eθ(H0)) / (1 − Eθ(H0)).
3.2 The HarnessOpt-Bench suite
The suite comprises OfficeQA, BrowseComp-Plus, Terminal-Bench, and GAIA, each with a deliberately untuned Python seed harness. Three seeds are competent but naive, whereas GAIA begins from a non-functional stub; seed and nominated-candidate test scores are pooled over K=3 rounds. Isolated sandboxes and a trusted gateway preserve candidate versions and enforce target-agent access and resource controls.
The diagram makes the held-out protocol operational: the optimizer can modify only target-agent code and can inspect permitted task data and feedback, while the evaluation server retains test data and budget enforcement. Isolated rollout sandboxes and gateway-mediated model calls reduce the opportunity to alter evaluation state or bypass resource controls.
4 Experimental Setup
Five optimizer models—claude-opus-5, claude-sonnet-5, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3—are evaluated under the shared opencode harness and their respective native harnesses, forming 10 core configurations. Held-out scores average three attempts per test case, core configurations are run twice, and task-specific resolution bands describe repeated-evaluation variation. LSS-λ estimates task-adjusted model effects using the balanced shared-harness grid over the three competent-seed tasks.
5 Results
Results are expressed as normalized gain relative to each task’s pinned seed, with task-specific resolution bands used to mark differences that remain unresolved at the measured repeatability scale. The analyses compare optimizer models, assess release-level progress, examine search behavior, and contrast shared with native optimizer harnesses.
The table gives mean normalized gain for each model–harness configuration on all four tasks and lists each task’s resolution band. It shows substantial task dependence and distinguishes GAIA, where the measured-zero seed makes gain equal to raw held-out score rather than a fraction of remaining headroom.
5.1 Distinguishing frontier models
Holding task and harness fixed, changing the optimizer model moves gain by 0.142 on average; holding task and model fixed, changing the harness moves it by 0.079, making the model contrast about 1.8× larger. On the shared-harness competent-seed scope, claude-opus-5 has LSS-λ +0.228 in tier 1, claude-sonnet-5, kimi-k3, and gpt-5.6-sol form an unresolved tier 2, and gpt-5.6-terra has −0.174 in tier 3.
5.2 Tracking model progress
On OfficeQA, five GPT releases increase monotonically from normalized gain +0.03 to +0.49, with three of four successive changes exceeding the ±0.045 resolution band. Five Claude Opus releases range from +0.37 to +0.59; the series is non-monotonic, but the earliest-to-latest difference exceeds that band.
With the target model, seed, budget, and coding harness held fixed within each release series, the plot tests whether the benchmark resolves model progress. GPT gains increase steadily across the displayed releases, while Claude Opus gains fluctuate despite an overall earliest-to-latest increase.
5.3 Where current optimizers fall short
Touching a larger fraction of eight pre-registered harness levers during search is positively associated with gain on every task, with Spearman ρ from +0.34 to +0.88; this is descriptive because lever breadth is correlated with modification volume and search effort. Detailed traces were requested only 16 times by 7 of 111 scored cells, and trace-reading share was negatively associated with gain (−0.31 to −0.64). Case-pass allowances, rather than evaluation-call caps, constrained search in the reported runs, and the best visible validation score generally exceeded the submitted candidate’s held-out score.
Each panel relates normalized gain to the fraction of pre-specified harness levers touched during search. The positive within-task Spearman correlations support an association between broader exploration and higher gain, but the figure does not establish causality because lever coverage is correlated with total modification volume.
5.4 The effect of the optimizer’s own harness
Across 20 paired model–task comparisons, opencode wins 11 times and the native harness wins 9 times, so native tooling is not uniformly superior. Effects are heterogeneous: on GAIA, codex exceeds the best shared harness by +0.179 for gpt-5.6-sol and +0.131 for gpt-5.6-terra, whereas Claude and Kimi native-harness differences lie within roughly one or two resolution bands.
The crossing lines show that optimizer-harness rankings on GAIA depend on the optimizer model. The native-versus-best-shared comparison isolates large codex advantages for the two GPT models, rather than evidence for a general advantage of native harnesses.
6 Conclusion
The authors conclude that harness optimization can be evaluated as a model capability, but gains remain uneven across tasks and often do not support a complete fine-grained ranking. Limitations include possible exploitation of stable development and validation evaluators, dependence on seed maturity, Python-only candidates, one pinned target model per task, and untested transfer across runtimes, architectures, and target models. The ethics statement identifies dual-use risks from automated agent improvement and describes bounded benign tasks, fixed model allow-lists, metered execution, and auditable versioning as safeguards.
Appendix
- Appendix A reports per-configuration gains, observed replicate ranges, explored lever coverage, and target-token use. Appendix B specifies task splits, pinned target models, seed baselines, and off-the-shelf harness scores; later appendices provide model effects, supporting figures, and reproducibility details covering immutable manifests, Git-versioned candidates, lockfiles, isolated sandboxes, and trace-derived process metrics.
Brief Thoughts
The benchmark’s central contribution is an enforced held-out evaluation boundary that makes claimed harness improvements auditable rather than dependent on optimizer self-report or visible validation scores. Its findings apply to the evaluated combination of seeds, tasks, target models, and uncapped optimizer inference, rather than constituting a comprehensive ranking of agent-improvement ability.