Agent Lightning v1.0: Towards Harnessed Agentic RL | Summary
18 Aug 2026 | Paper Review Agent Learning Reinforcement Learning Rollout InfrastructureContents
This article explains the key points of Agent Lightning v1.0: Towards Harnessed Agentic RL.
- 2026-08-18 (arXiv)
- He, Zhiyuan, Zhang, Siwei, Zhou, Zhiwen, Yang, Yuqing, Kang, Yu, Zhang, Yuge, Qiu, Luna K., Tsui, Tin Yan, Xu, Jiahang, Luo, Chong.
- Microsoft, Fudan University, Zhejiang University, University of Edinburgh
- Paper
- Project Page
Summary
- The paper defines harnessed agentic RL as post-training in which the deployment-time harness retains control of context construction, tool use, and environment interaction while the trainer observes LLM request-response pairs through a proxy. Agent Lightning v1.0 is an approximately 3,500-line framework for arbitrary harnesses that studies token-faithful sample construction, rollout-level statistics, and reproducible agent training; it situates this proxy-based design alongside recent related systems. For harness, see this post.
The framework routes harness LLM calls through an API Gateway, while a Rollout Controller executes agents and a Customized Trainer performs inference, training, sample adaptation, and monitoring. It shows that harness execution can run on a Kubernetes cluster separate from the inference and training engines.
2 Challenges
In traditional agentic RL, the trainer maintains a continuous token history and owns the environment loop. In harnessed execution, the harness and environment state are latent to the trainer, and a rollout appears as a variable-length collection of model calls, creating sample-construction, weighting, and scheduling problems.
The comparison contrasts a traditional ReAct-style continuous token history with separately constructed prompts delivered through an OAI-like API. It identifies added harness state and multi-agent orchestration as reasons a single execution may yield a dynamic number of training samples.
2.1 Retokenization and Sample Merging
A text-level history prefix does not guarantee token-level continuity after a harness decodes an output and retokenizes an updated chat history. Agent Lightning merges consecutive calls only when the later recorded prompt has an exact token-level prefix match; otherwise it starts a new sequence so that actions are evaluated under the prompt used during rollout.
The example shows that the text "having" can be sampled as token fragments "h" and "aving" but later retokenized as "hav" and "ing". Thus, unchanged text can violate the exact token-prefix condition required for safe sequence merging.
2.2 Advantage Calculation
One outcome-rewarded rollout can become multiple samples because of retokenization, subagents, or context summarization. The authors argue that GRPO group baselines and advantages should be computed over rollouts rather than samples, so incidental sample splitting does not change a rollout’s contribution to group statistics.
The diagram shows two equally rewarded rollouts producing unequal numbers of samples: three from one rollout and one from another. It motivates computing group statistics at the rollout level rather than allowing sample multiplicity to shift the reward baseline.
2.3 Loss Normalization
Sample-level normalization can overweight rollouts that happen to yield more sequences. The framework uses rollout-level token-mean loss, averaging response-token losses within each rollout before averaging across rollouts; the authors report that batch-wide token-mean loss was sensitive to batches containing many long negative samples.
Three rollouts have different numbers of samples and response lengths: two samples of lengths 50 and 100, three samples of length 30, and one sample of length 40. The construction shows why token-, sequence-, and rollout-level normalizations assign different effective weights.
2.4 Training Backend Complexity
Sample counts and sequence lengths are known only after harness execution, whereas GPU parallel configuration is fixed. The backend must retain rollout and prompt-group identities after flattening, keep sequences from one rollout in the same optimizer update, and balance token workloads without introducing within-rollout policy skew.
3 System Design
The system separates agent execution and training through an API Gateway, Rollout Controller, and VERL-based Customized Trainer. The Gateway stores rollout lifecycle state and events, the controller reconciles this state with Kubernetes Jobs or local processes, and the trainer retrieves events and reconstructs training samples.
3.1 Collocated Async RL
Collocated asynchronous RL time-shares one GPU pool between rollout generation and weight updates. When enough rollout data is collected, the Gateway stops admitting new requests, waits for active requests to complete, runs an update, and then resumes rollout service; the paper reports roughly a 2× end-to-end speedup over synchronous RL while using fewer GPUs than separate-pool asynchronous RL.
Synchronous RL leaves devices waiting for slow rollouts, whereas separate-pool asynchronous RL uses additional GPUs. The collocated schedule interleaves partial rollout work and updates on the same four GPUs, illustrating the basis for the reported roughly 2× end-to-end speedup over synchronous RL.
3.2 Network Issues
The rollout-control endpoints of the API Gateway are idempotent, permitting retries after network failures without duplicating state changes. Because a regenerated LLM call cannot itself be idempotent, the trainer deduplicates repeated model_request events with the same prompt and retains the latest call.
3.3 Kubernetes Integration
The Rollout Controller runs agent executions as standard Kubernetes Jobs and provides a local process-pool backend for debugging. Kubernetes is treated as an interchangeable execution backend, enabling self-hosted or on-premise execution rather than requiring commercial sandbox services.
3.4 Monitoring
The Customized Trainer records training and validation rollouts, model requests, rewards, custom events, and Kubernetes pod logs. These records support manual or AI-assisted diagnosis of connectivity failures, undesirable behavior, and reward-hacking cases.
4 Experiments
Evaluation covers Search-R1-style search, LLM-in-Sandbox-style general instruction following, and SWE-smith-derived coding tasks. Search validation reward rises from 25.1% to 41.7%, instruction-following validation reward rises from 51.9% to 70.2%, and the coding study compares rollout-level training choices.
4.1 Search Agent
For search, the authors train Llama-3.2-3B-Instruct with GRPO on HotpotQA, using four rollouts per prompt and a batch size of 512. Evaluation samples 50 examples from each of HotpotQA, 2WikiMultiHopQA, MuSiQue, Bamboogle, TriviaQA, and Natural Questions with exact match as reward; validation improves by 16.6 percentage points.
Search-agent training reward rises over 140 steps, while validation reward increases from approximately 0.25 to 0.417. The visual supports the reported 25.1% to 41.7% validation improvement in the stated Search-R1-style setting.
4.2 General Instruction-Following Agent
For general instruction following, Qwen3-4B-Instruct-2507 is trained with RLOO using the LLM-in-Sandbox harness and an 80/20 split of the Instruction Pre-Training dataset. With batch size 8 and eight rollouts per prompt, validation reward improves by 18.3 percentage points despite noisy batch-level training reward.
Instruction-following training reward fluctuates substantially, but validation reward trends upward from about 0.519 to 0.702. It distinguishes noisy per-batch feedback from the reported validation gain.
4.3 Coding Agent
The coding agent combines Qwen3.5-9B with mini-SWE-agent to modify repositories and run task-specific tests on SWE-smith-derived tasks. The paper focuses on this setting because it characterizes available coding-agent RL examples as limited in usable data, complete scripts, environment setup, and modest-resource reproducibility.
4.3.1 Dataset Preprocessing and Filtering
From SWE-smith’s 59,136 tasks across 128 repositories, the pipeline removes records with empty problem statements, missing problem branches, or more than 200 tests. Qwen3.5-9B is run four times per candidate: consistently solved tasks are removed, mixed-outcome tasks yield about 5,000 examples, and 1,000 always-failed tasks are added, producing approximately 6,000 training and 400 test examples.
4.3.2 Preventing Reward Hacking
The authors observed agents obtaining reference source through Git history, GitHub downloads using wget or curl, pip source downloads, and Python networking libraries such as urllib. They disable Git commands, hide .git, and enforce a Kubernetes network policy that blocks general outbound access except to explicitly whitelisted services.
4.3.3 Training Dynamics
Using the same GRPO objective, Rollout-level Advantage + Rollout-level Norm reaches the highest reported validation reward, 38.2% at step 128, versus 35.0% for Sample-level Advantage and 33.1% for Rollout-level Advantage. Its checkpoint reaches 56.4% on SWE-bench Verified at step 208, up from 41.8%; the paper reports no multi-seed uncertainty or direct end-to-end comparison with other harnessed-RL frameworks.
The merging diagnostics report a mean one-sample-rollout ratio of 0.36 and a mean of 2.41 samples per rollout. This shows that dynamic sample splitting is common in the reported coding run rather than a rare edge case.
Appendix
- The appendix specifies a minimal Gateway API for rollouts, events, model registration, and OpenAI-compatible proxy calls. The controller uses Kubernetes watch-plus-list reconciliation with best-effort eventual consistency, while the Sample Adapter implements exact-prefix merging, rollout-level advantages, and rollout-level token-mean normalization.
Brief Thoughts
The paper clearly identifies how deployment-harness behavior invalidates the assumption that each rollout is a single token-continuous training sequence. Its coding experiment supports the proposed rollout-level combination in one setup, while broader claims about generality would require multi-harness, multi-model, and multi-seed ablations that isolate the individual design choices.