Gorio Tech Blog search

EnvHarness: Awakening Static Worlds for Agent Learning | Summary

|

Contents

This article explains the key points of EnvHarness: Awakening Static Worlds for Agent Learning.

  • 2026-08-20 (arXiv)
  • Huang, Chengsong, Wang, Zifeng, Han, Rujun, Yan, Jun, Chen, Yanfei, CuiZhu, Zoey, Jiang, Ke, Xia, Peng, Yu, Han, Zhuang, Yufan, et al.
  • Washington University in St. Louis, Google Cloud AI Research, Google Cloud, University of North Carolina at Chapel Hill
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • EnvHarness treats environment construction for LLM agents as a wrapping problem: plug-in components reshape a static environment through its standard interface without modifying its underlying implementation or replacing its verifier. EnvRigger diagnoses a black-box policy from rollouts, writes targeted wrappers, and validates them on fresh rollouts; across five benchmarks in four domains, the paper reports up to a 9.0-point held-out improvement and 9.8% fewer execution steps.

The left bars compare a base agent, agents learning from original environments, and agents learning from EnvHarness environments on SWE-bench Verified, OfficeQA, and SpreadsheetBench; the EnvHarness condition is highest in all three groups. The right plot shows its resolved rate increasing through 300 environments, whereas the original-environment and generated-environment curves flatten at lower values.

Overall performance across software-engineering and office-automation benchmarks, together with SWE-bench Verified scaling under an identical environment budget.
Overall performance across software-engineering and office-automation benchmarks, together with SWE-bench Verified scaling under an identical environment budget.

1. Introduction

The paper argues that hand-built environments are static with respect to an improving agent: they neither target current weaknesses nor continue to provide useful learning signals after tasks become easy. It presents existing environment-generation pipelines as domain-specific and difficult to verify, whereas EnvHarness mediates reset/step interaction externally while retaining the original task and verifier.

2. EnvHarness

EnvHarness is a programmable layer around an existing environment that changes information flow without editing the simulator backend. Its plug-in wrappers can alter initial states, interaction constraints, observations, and episode composition while preserving the standard environment interface for the agent.

2.1. The EnvHarness Paradigm

The paper models an environment as E = (S, A, O, T, R, s0) and defines a component as an environment-agnostic transformation w that produces E′ = w(E). The wrapper may alter exposed state, action, observation, and transition behavior at the interface boundary while the original evaluator remains available to score the episode.

2.2. Three EnvHarness Components

Stage changes an episode’s initial state by replaying valid environment actions after reset. Contract rewrites actions, transitions, or observations through maps (fA, fT, fO), while Chain combines environments into composite episodes; component composition is non-commutative, so wrapper order determines the resulting initialization and interaction constraints.

The figure locates each wrapper at a standard-interface intervention point: Stage modifies reset-derived initialization, Contract intercepts step-time action, transition, or observation handling, and Chain changes episode-level composition. It emphasizes that native state transitions and the original verifier remain in the frozen base environment.

Overview of Stage, Contract, and Chain wrappers around the standard environment interface.
Overview of Stage, Contract, and Chain wrappers around the standard environment interface.

3. EnvHarness for Agent Learning

For agent learning, the framework seeks a task- and policy-conditioned wrapped environment that exposes a target policy’s specific weaknesses. Individual components are policy-agnostic transformations, but their selection and parameterization are conditioned on the base task and observed policy behavior.

3.1. Problem Setup

Given a base environment E, task t, and target policy π, the objective is to generate E′ = H(E, t; π) by composing wrappers that provide a corrective learning signal. The method treats π as a black box, using outputs and trajectories rather than model weights to identify behavioral vulnerabilities.

3.2. EnvRigger

EnvRigger has four stages: Observe policy rollouts, Diagnose systematic failures, Write one or more candidate components, and Validate candidates on fresh rollouts. Candidates can be accepted, rejected as trivial or unsolvable, or refined when their learning signal is poorly scaled; the implementation uses 5 baseline rollouts, 5 validation rollouts, and at most 5 write–validate rounds per instance.

The workflow separates target-policy execution in the currently wrapped environment from EnvRigger’s Observe, Diagnose, Write, and Validate loop. Fresh rollout evidence determines whether candidate components are accepted, rejected, or revised before accepted components are appended to the active wrapper stack.

EnvRigger workflow for generating and validating EnvHarness components from a target policy’s trajectories.
EnvRigger workflow for generating and validating EnvHarness components from a target policy’s trajectories.

4. Experiments

The primary skill-based experiments cover ALFWorld, WebArena, SWE-bench Verified, OfficeQA, and SpreadsheetBench, with disjoint training and evaluation episodes. EnvRigger and the policy use the same backbone for each benchmark—Gemini-3.1-Flash-Lite for ALFWorld and WebArena, and Gemini-3.5-Flash elsewhere—and skills are extracted following ReasoningBank; Chain is excluded from this automated pipeline and studied separately.

4.1. Experimental Setup

The baselines are No Skills, skills from Original Envs, and the benchmark-specific generation methods GenEnv, VeriEnv, or SWE-smith where applicable. The authors hold seed instances, environment count, skill-extraction pipeline, and policy model fixed; training uses 100 ALFWorld tasks, 20 WebArena tasks per sub-domain, 100 SWE-bench Lite tasks, 50 OfficeQA tasks, and 100 SpreadsheetBench tasks, followed by evaluation on held-out original tasks.

4.2. Main Results

EnvHarness-derived skills exceed Original Envs on every reported benchmark. The largest margin is +9.0 points on ALFWorld OOD; on SWE-bench Verified, success rate is 52.58 versus 49.88 and average steps are 49.61 versus 55.01, while OfficeQA and SpreadsheetBench also improve on their displayed metrics. Table 2 reports means over three independent runs, with standard deviations shown for Tables 2 and 3.

On ALFWorld, EnvHarness reaches 68.3 average performance versus 62.4 for Original Envs, including 70.4 OOD versus 61.4. On WebArena, it reaches 41.6 average versus 38.5 for Original Envs and 39.6 for VeriEnv, with its largest displayed gain in Shop Admin.

Skill-source comparison on ALFWorld and WebArena, reporting mean performance over three independent runs.
Skill-source comparison on ALFWorld and WebArena, reporting mean performance over three independent runs.

EnvHarness yields 52.58 SWE-bench Verified success rate with 49.61 average steps, compared with 49.88 and 55.01 for Original Envs. It also has the highest displayed OfficeQA and SpreadsheetBench metrics, while Original Envs slightly underperform No Skills on SpreadsheetBench Pass@1.

Skill-source comparison on SWE-bench Verified, OfficeQA, and SpreadsheetBench.
Skill-source comparison on SWE-bench Verified, OfficeQA, and SpreadsheetBench.

5. Analysis

In RL, Qwen3-8B-base trained with GRPO on EnvHarness environments improves ALFWorld in-distribution success from 81.4 to 87.9 and WebShop score/success rate from 75.6/66.0 to 79.2/67.4, while ALFWorld OOD success decreases from 89.6 to 88.8. For long-horizon SWE-bench analysis, Chain-only skills reduce average steps to 41.96 but attain 49.63 success rate; combining Chain with Stage/Contract skills yields 54.30 success rate and 43.12 average steps.

The RL comparison shows EnvHarness-trained policies outperforming original-environment-trained policies on ALFWorld in-distribution success and on both WebShop metrics. The exception is a 0.8-point decrease in ALFWorld OOD success, so the reported gains are not uniform across outcomes.

Reinforcement-learning results for policies trained on original versus EnvHarness environments in ALFWorld and WebShop.
Reinforcement-learning results for policies trained on original versus EnvHarness environments in ALFWorld and WebShop.

Chain-only skills reduce average steps from 53.58 for No Skills to 41.96, but their 49.63 success rate is below the 49.88 reported for Original Envs. Combining Chain with Stage/Contract skills gives the highest displayed success rate, 54.30, with a 43.12 average-step count.

Long-horizon SWE-bench results isolating Stage/Contract skills, Chain skills, and their combination.
Long-horizon SWE-bench results isolating Stage/Contract skills, Chain skills, and their combination.

With identical environment counts and a shared extraction and retrieval protocol, EnvHarness rises from 47.67 to 54.79 resolved rate at 300 environments. Original environments reach 52.13 and generated environments 50.37, providing the paper’s scaling evidence for iterative policy-conditioned environment construction.

SWE-bench Verified resolved rate as the number of original, generated, or EnvHarness environments increases.
SWE-bench Verified resolved rate as the number of original, generated, or EnvHarness environments increases.

Across Gemini 3.1 Flash-Lite, Qwen3.6 27B, Gemini 3.5 Flash, and Claude Sonnet 4.6, EnvHarness skill banks have the highest success-rate bar in every group. The absolute advantage over Original Envs ranges from 2.7 to 3.7 points.

Cross-model SWE-bench Verified success rates with no skills, skills from original environments, and skills from EnvHarness environments.
Cross-model SWE-bench Verified success rates with no skills, skills from original environments, and skills from EnvHarness environments.

The related-work discussion distinguishes EnvHarness from simulators and programmatic environment-synthesis systems by emphasizing interface-level adaptation of existing environments rather than construction of new simulators, tools, tasks, or verifiers. It also contrasts the method with self-evolving-agent work by adapting the environment presented to a fixed policy in response to rollout-based diagnosis.

6.1. Environment Scaling

Environment-scaling methods may simulate feedback, synthesize executable environments, generate new task instances, or design curricula. EnvHarness instead uses a common wrapper interface across benchmarks, conditions customization on diagnosed policy weaknesses, and retains the tasks and verifiers of the wrapped environments.

6.2. Self-Evolving Agent

Prior self-evolving-agent methods update prompts, reflections, skills, memories, workflows, harnesses, model weights, rewards, or self-proposed tasks. EnvHarness positions its contribution as adapting the training environment while treating the target policy as a black-box object during diagnosis.

7. Conclusion

The conclusion presents Stage, Contract, and Chain as reset/step wrappers for difficulty calibration, skill isolation, and horizon extension without modifying base environment code. Its central practical claim is that inherited verification can be combined with EnvRigger’s rollout-driven component synthesis.

Appendix

  • The appendix specifies the EnvRigger prompt, ActionableEnv protocol, Bridges for heterogeneous runtimes, decorator-based component stacks, Chain implementations, benchmark splits, baselines, hyperparameters, and supplementary analyses. It notes that the iterative design loop can be costly; Stage and Chain require a resettable Gym-style interface; and the implemented Chain experiments use serial composition, while semantically grounded branching or shared intermediate state would require a composite verifier. Leave-one-out ALFWorld transfer improves by 3.1 points on average but regresses by 8.7 points on the heat task.

Brief Thoughts

The controlled comparisons use shared policy models, extraction procedures, seed instances, and environment counts within each benchmark, and the scaling experiment evaluates repeated policy-conditioned environment construction under a fixed budget. The evidence is nevertheless limited to resettable, predominantly text-based settings, small rollout-based validation samples, and wrappers that may create synthetic constraints; transfer to untouched tasks is positive on average in the reported leave-one-out analysis but not uniform.