Gorio Tech Blog search

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization | Summary

|

Contents

This article explains the key points of HarnessOpt-Bench: Evaluating LLMs at Harness Optimization.

  • 2026-08-06 (arXiv)
  • Ursekar, Varun, Shanker, Apaar, Maurya, Yash, Yasser, Shehab, Kalmath, Vijay S., Chatrath, Veronica, Xue, Yuan.
  • Scale AI
  • Paper

Read this article in Korean


Summary

  • HARNESSOPT-BENCH evaluates whether an LLM operating through a coding harness can improve a target agent’s prompts, tools, control flow, memory, and orchestration code under expensive stochastic evaluation. Across 111 scored runs on four downstream tasks, optimizer-model differences are larger on average than coding-harness differences, while native coding harnesses show no consistent advantage.

1 Introduction

The paper treats harness optimization as harder than ordinary code editing because the optimizer must infer improvement from noisy and costly agent evaluations under a fixed target-evaluation budget. It holds the target model, environment, and verifier fixed; reserves final scoring for an inaccessible test partition; and uses a trusted boundary to enforce access restrictions, resource accounting, and candidate auditability.

Related work includes prompt and textual-artifact optimization, evolutionary program search, and code-defined agent-design search. The closest methods use coding agents as end-to-end harness optimizers, but their reported results often vary jointly with seeds, budgets, search spaces, and evaluation protocols; HARNESSOPT-BENCH fixes these factors to support controlled comparison.

3 HarnessOpt-Bench: Harness Optimization as a Task

The benchmark defines harness optimization as constrained stochastic program optimization. An optimizer edits a pinned executable seed harness, evaluates candidate versions through graded feedback within a resource budget, and nominates one candidate for hidden test evaluation.

3.1 The optimization problem

A task fixes invariants θ = (M, E, V): the target models available to candidates, each case’s environment, and a verifier that maps trajectories to scores in [0, 1]. The optimizer may modify the feasible harness H but not θ; development feedback reveals inputs, per-case outcomes, and traces, validation reveals aggregate scores, and test feedback remains unavailable until nomination. The benchmark reports normalized gain g = (Eθ(H+) − Eθ(H0)) / (1 − Eθ(H0)).

3.2 The HarnessOpt-Bench suite

The suite comprises OfficeQA, BrowseComp-Plus, Terminal-Bench, and GAIA, each with a deliberately untuned Python seed harness. Three seeds are competent but naive, whereas GAIA begins from a non-functional stub; seed and nominated-candidate test scores are pooled over K=3 rounds. Isolated sandboxes and a trusted gateway preserve candidate versions and enforce target-agent access and resource controls.

The diagram makes the held-out protocol operational: the optimizer can modify only target-agent code and can inspect permitted task data and feedback, while the evaluation server retains test data and budget enforcement. Isolated rollout sandboxes and gateway-mediated model calls reduce the opportunity to alter evaluation state or bypass resource controls.

Trusted execution architecture separating optimizer access, target-harness editing, evaluation sandboxes, model calls, and hidden test evaluation.
Trusted execution architecture separating optimizer access, target-harness editing, evaluation sandboxes, model calls, and hidden test evaluation.

4 Experimental Setup

Five optimizer models—claude-opus-5, claude-sonnet-5, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3—are evaluated under the shared opencode harness and their respective native harnesses, forming 10 core configurations. Held-out scores average three attempts per test case, core configurations are run twice, and task-specific resolution bands describe repeated-evaluation variation. LSS-λ estimates task-adjusted model effects using the balanced shared-harness grid over the three competent-seed tasks.

5 Results

Results are expressed as normalized gain relative to each task’s pinned seed, with task-specific resolution bands used to mark differences that remain unresolved at the measured repeatability scale. The analyses compare optimizer models, assess release-level progress, examine search behavior, and contrast shared with native optimizer harnesses.

The table gives mean normalized gain for each model–harness configuration on all four tasks and lists each task’s resolution band. It shows substantial task dependence and distinguishes GAIA, where the measured-zero seed makes gain equal to raw held-out score rather than a fraction of remaining headroom.

Normalized gain by optimizer configuration and downstream task, with task-specific resolution bands.
Normalized gain by optimizer configuration and downstream task, with task-specific resolution bands.

5.1 Distinguishing frontier models

Holding task and harness fixed, changing the optimizer model moves gain by 0.142 on average; holding task and model fixed, changing the harness moves it by 0.079, making the model contrast about 1.8× larger. On the shared-harness competent-seed scope, claude-opus-5 has LSS-λ +0.228 in tier 1, claude-sonnet-5, kimi-k3, and gpt-5.6-sol form an unresolved tier 2, and gpt-5.6-terra has −0.174 in tier 3.

5.2 Tracking model progress

On OfficeQA, five GPT releases increase monotonically from normalized gain +0.03 to +0.49, with three of four successive changes exceeding the ±0.045 resolution band. Five Claude Opus releases range from +0.37 to +0.59; the series is non-monotonic, but the earliest-to-latest difference exceeds that band.

With the target model, seed, budget, and coding harness held fixed within each release series, the plot tests whether the benchmark resolves model progress. GPT gains increase steadily across the displayed releases, while Claude Opus gains fluctuate despite an overall earliest-to-latest increase.

OfficeQA normalized gain across successive Claude Opus and GPT optimizer-model releases.
OfficeQA normalized gain across successive Claude Opus and GPT optimizer-model releases.

5.3 Where current optimizers fall short

Touching a larger fraction of eight pre-registered harness levers during search is positively associated with gain on every task, with Spearman ρ from +0.34 to +0.88; this is descriptive because lever breadth is correlated with modification volume and search effort. Detailed traces were requested only 16 times by 7 of 111 scored cells, and trace-reading share was negatively associated with gain (−0.31 to −0.64). Case-pass allowances, rather than evaluation-call caps, constrained search in the reported runs, and the best visible validation score generally exceeded the submitted candidate’s held-out score.

Each panel relates normalized gain to the fraction of pre-specified harness levers touched during search. The positive within-task Spearman correlations support an association between broader exploration and higher gain, but the figure does not establish causality because lever coverage is correlated with total modification volume.

Relationship between explored harness-lever breadth and normalized gain across four tasks.
Relationship between explored harness-lever breadth and normalized gain across four tasks.

5.4 The effect of the optimizer’s own harness

Across 20 paired model–task comparisons, opencode wins 11 times and the native harness wins 9 times, so native tooling is not uniformly superior. Effects are heterogeneous: on GAIA, codex exceeds the best shared harness by +0.179 for gpt-5.6-sol and +0.131 for gpt-5.6-terra, whereas Claude and Kimi native-harness differences lie within roughly one or two resolution bands.

The crossing lines show that optimizer-harness rankings on GAIA depend on the optimizer model. The native-versus-best-shared comparison isolates large codex advantages for the two GPT models, rather than evidence for a general advantage of native harnesses.

GAIA optimizer-harness sweep and each model’s native-harness difference from its best shared-harness result.
GAIA optimizer-harness sweep and each model’s native-harness difference from its best shared-harness result.

6 Conclusion

The authors conclude that harness optimization can be evaluated as a model capability, but gains remain uneven across tasks and often do not support a complete fine-grained ranking. Limitations include possible exploitation of stable development and validation evaluators, dependence on seed maturity, Python-only candidates, one pinned target model per task, and untested transfer across runtimes, architectures, and target models. The ethics statement identifies dual-use risks from automated agent improvement and describes bounded benign tasks, fixed model allow-lists, metered execution, and auditable versioning as safeguards.

Appendix

  • Appendix A reports per-configuration gains, observed replicate ranges, explored lever coverage, and target-token use. Appendix B specifies task splits, pinned target models, seed baselines, and off-the-shelf harness scores; later appendices provide model effects, supporting figures, and reproducibility details covering immutable manifests, Git-versioned candidates, lockfiles, isolated sandboxes, and trace-derived process metrics.

Brief Thoughts

The benchmark’s central contribution is an enforced held-out evaluation boundary that makes claimed harness improvements auditable rather than dependent on optimizer self-report or visible validation scores. Its findings apply to the evaluated combination of seeds, tasks, target models, and uncapped optimizer inference, rather than constituting a comprehensive ranking of agent-improvement ability.