Gorio Tech Blog search

Agent Lightning v1.0: Towards Harnessed Agentic RL | Summary

|

Contents

This article explains the key points of Agent Lightning v1.0: Towards Harnessed Agentic RL.

  • 2026-08-18 (arXiv)
  • He, Zhiyuan, Zhang, Siwei, Zhou, Zhiwen, Yang, Yuqing, Kang, Yu, Zhang, Yuge, Qiu, Luna K., Tsui, Tin Yan, Xu, Jiahang, Luo, Chong.
  • Microsoft, Fudan University, Zhejiang University, University of Edinburgh
  • Paper
  • Project Page

Read this article in Korean


Summary

  • The paper defines harnessed agentic RL as post-training in which the deployment-time harness retains control of context construction, tool use, and environment interaction while the trainer observes LLM request-response pairs through a proxy. Agent Lightning v1.0 is an approximately 3,500-line framework for arbitrary harnesses that studies token-faithful sample construction, rollout-level statistics, and reproducible agent training; it situates this proxy-based design alongside recent related systems. For harness, see this post.

The framework routes harness LLM calls through an API Gateway, while a Rollout Controller executes agents and a Customized Trainer performs inference, training, sample adaptation, and monitoring. It shows that harness execution can run on a Kubernetes cluster separate from the inference and training engines.

Overall architecture connecting harnessed agents, the Gateway, rollout control, inference, and training.
Overall architecture connecting harnessed agents, the Gateway, rollout control, inference, and training.

2 Challenges

In traditional agentic RL, the trainer maintains a continuous token history and owns the environment loop. In harnessed execution, the harness and environment state are latent to the trainer, and a rollout appears as a variable-length collection of model calls, creating sample-construction, weighting, and scheduling problems.

The comparison contrasts a traditional ReAct-style continuous token history with separately constructed prompts delivered through an OAI-like API. It identifies added harness state and multi-agent orchestration as reasons a single execution may yield a dynamic number of training samples.

Traditional agentic RL versus harnessed agentic RL in state, model input, and agent structure.
Traditional agentic RL versus harnessed agentic RL in state, model input, and agent structure.

2.1 Retokenization and Sample Merging

A text-level history prefix does not guarantee token-level continuity after a harness decodes an output and retokenizes an updated chat history. Agent Lightning merges consecutive calls only when the later recorded prompt has an exact token-level prefix match; otherwise it starts a new sequence so that actions are evaluated under the prompt used during rollout.

The example shows that the text "having" can be sampled as token fragments "h" and "aving" but later retokenized as "hav" and "ing". Thus, unchanged text can violate the exact token-prefix condition required for safe sequence merging.

Retokenization can change token boundaries even when the text is unchanged.
Retokenization can change token boundaries even when the text is unchanged.

2.2 Advantage Calculation

One outcome-rewarded rollout can become multiple samples because of retokenization, subagents, or context summarization. The authors argue that GRPO group baselines and advantages should be computed over rollouts rather than samples, so incidental sample splitting does not change a rollout’s contribution to group statistics.

The diagram shows two equally rewarded rollouts producing unequal numbers of samples: three from one rollout and one from another. It motivates computing group statistics at the rollout level rather than allowing sample multiplicity to shift the reward baseline.

A traditional one-rollout-one-sample setting compared with dynamic sample expansion under a harness.
A traditional one-rollout-one-sample setting compared with dynamic sample expansion under a harness.

2.3 Loss Normalization

Sample-level normalization can overweight rollouts that happen to yield more sequences. The framework uses rollout-level token-mean loss, averaging response-token losses within each rollout before averaging across rollouts; the authors report that batch-wide token-mean loss was sensitive to batches containing many long negative samples.

Three rollouts have different numbers of samples and response lengths: two samples of lengths 50 and 100, three samples of length 30, and one sample of length 40. The construction shows why token-, sequence-, and rollout-level normalizations assign different effective weights.

Batch example with unequal rollout sample counts and response lengths.
Batch example with unequal rollout sample counts and response lengths.

2.4 Training Backend Complexity

Sample counts and sequence lengths are known only after harness execution, whereas GPU parallel configuration is fixed. The backend must retain rollout and prompt-group identities after flattening, keep sequences from one rollout in the same optimizer update, and balance token workloads without introducing within-rollout policy skew.

3 System Design

The system separates agent execution and training through an API Gateway, Rollout Controller, and VERL-based Customized Trainer. The Gateway stores rollout lifecycle state and events, the controller reconciles this state with Kubernetes Jobs or local processes, and the trainer retrieves events and reconstructs training samples.

3.1 Collocated Async RL

Collocated asynchronous RL time-shares one GPU pool between rollout generation and weight updates. When enough rollout data is collected, the Gateway stops admitting new requests, waits for active requests to complete, runs an update, and then resumes rollout service; the paper reports roughly a 2× end-to-end speedup over synchronous RL while using fewer GPUs than separate-pool asynchronous RL.

Synchronous RL leaves devices waiting for slow rollouts, whereas separate-pool asynchronous RL uses additional GPUs. The collocated schedule interleaves partial rollout work and updates on the same four GPUs, illustrating the basis for the reported roughly 2× end-to-end speedup over synchronous RL.

Scheduling comparison among synchronous, asynchronous, and collocated asynchronous RL.
Scheduling comparison among synchronous, asynchronous, and collocated asynchronous RL.

3.2 Network Issues

The rollout-control endpoints of the API Gateway are idempotent, permitting retries after network failures without duplicating state changes. Because a regenerated LLM call cannot itself be idempotent, the trainer deduplicates repeated model_request events with the same prompt and retains the latest call.

3.3 Kubernetes Integration

The Rollout Controller runs agent executions as standard Kubernetes Jobs and provides a local process-pool backend for debugging. Kubernetes is treated as an interchangeable execution backend, enabling self-hosted or on-premise execution rather than requiring commercial sandbox services.

3.4 Monitoring

The Customized Trainer records training and validation rollouts, model requests, rewards, custom events, and Kubernetes pod logs. These records support manual or AI-assisted diagnosis of connectivity failures, undesirable behavior, and reward-hacking cases.

4 Experiments

Evaluation covers Search-R1-style search, LLM-in-Sandbox-style general instruction following, and SWE-smith-derived coding tasks. Search validation reward rises from 25.1% to 41.7%, instruction-following validation reward rises from 51.9% to 70.2%, and the coding study compares rollout-level training choices.

4.1 Search Agent

For search, the authors train Llama-3.2-3B-Instruct with GRPO on HotpotQA, using four rollouts per prompt and a batch size of 512. Evaluation samples 50 examples from each of HotpotQA, 2WikiMultiHopQA, MuSiQue, Bamboogle, TriviaQA, and Natural Questions with exact match as reward; validation improves by 16.6 percentage points.

Search-agent training reward rises over 140 steps, while validation reward increases from approximately 0.25 to 0.417. The visual supports the reported 25.1% to 41.7% validation improvement in the stated Search-R1-style setting.

Training and validation reward trajectories for the search agent.
Training and validation reward trajectories for the search agent.

4.2 General Instruction-Following Agent

For general instruction following, Qwen3-4B-Instruct-2507 is trained with RLOO using the LLM-in-Sandbox harness and an 80/20 split of the Instruction Pre-Training dataset. With batch size 8 and eight rollouts per prompt, validation reward improves by 18.3 percentage points despite noisy batch-level training reward.

Instruction-following training reward fluctuates substantially, but validation reward trends upward from about 0.519 to 0.702. It distinguishes noisy per-batch feedback from the reported validation gain.

Training and validation reward trajectories for the general instruction-following agent.
Training and validation reward trajectories for the general instruction-following agent.

4.3 Coding Agent

The coding agent combines Qwen3.5-9B with mini-SWE-agent to modify repositories and run task-specific tests on SWE-smith-derived tasks. The paper focuses on this setting because it characterizes available coding-agent RL examples as limited in usable data, complete scripts, environment setup, and modest-resource reproducibility.

4.3.1 Dataset Preprocessing and Filtering

From SWE-smith’s 59,136 tasks across 128 repositories, the pipeline removes records with empty problem statements, missing problem branches, or more than 200 tests. Qwen3.5-9B is run four times per candidate: consistently solved tasks are removed, mixed-outcome tasks yield about 5,000 examples, and 1,000 always-failed tasks are added, producing approximately 6,000 training and 400 test examples.

4.3.2 Preventing Reward Hacking

The authors observed agents obtaining reference source through Git history, GitHub downloads using wget or curl, pip source downloads, and Python networking libraries such as urllib. They disable Git commands, hide .git, and enforce a Kubernetes network policy that blocks general outbound access except to explicitly whitelisted services.

4.3.3 Training Dynamics

Using the same GRPO objective, Rollout-level Advantage + Rollout-level Norm reaches the highest reported validation reward, 38.2% at step 128, versus 35.0% for Sample-level Advantage and 33.1% for Rollout-level Advantage. Its checkpoint reaches 56.4% on SWE-bench Verified at step 208, up from 41.8%; the paper reports no multi-seed uncertainty or direct end-to-end comparison with other harnessed-RL frameworks.

The merging diagnostics report a mean one-sample-rollout ratio of 0.36 and a mean of 2.41 samples per rollout. This shows that dynamic sample splitting is common in the reported coding run rather than a rare edge case.

Frequency of one-sample rollouts and mean samples per rollout during coding-agent training.
Frequency of one-sample rollouts and mean samples per rollout during coding-agent training.

Appendix

  • The appendix specifies a minimal Gateway API for rollouts, events, model registration, and OpenAI-compatible proxy calls. The controller uses Kubernetes watch-plus-list reconciliation with best-effort eventual consistency, while the Sample Adapter implements exact-prefix merging, rollout-level advantages, and rollout-level token-mean normalization.

Brief Thoughts

The paper clearly identifies how deployment-harness behavior invalidates the assumption that each rollout is a single token-continuous training sequence. Its coding experiment supports the proposed rollout-level combination in one setup, while broader claims about generality would require multi-harness, multi-model, and multi-seed ablations that isolate the individual design choices.