PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost | Summary
22 Mar 2026 | Paper Review Agent Learning Reinforcement Learning Tool UseContents
- Summary
- 1. Introduction
- 2. Preliminaries and Motivating Observations
- 3. PivotRL
- 4. Experiments
- 5. Related Works
- 6. Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost.
- 2026-03-22 (arXiv)
- Yi, Junkeun, Mosk-Aoyama, Damon, Huang, Baihe, Gala, Ritu, Wang, Charles, Devare, Sugam Dipak, Bhardwaj, Khushi, Gupta, Abhibha, Kuchaiev, Oleksii, Jiao, Jiantao, et al.
- NVIDIA, UC Berkeley
- Paper
Summary
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost converts expert demonstrations into local, on-policy reinforcement learning episodes for long-horizon agentic tasks. It selects intermediate assistant turns with mixed success/failure outcomes and rewards locally acceptable actions rather than requiring exact reproduction of the demonstrated completion.
- Across four separately trained agentic domains, PivotRL improves average in-domain accuracy by +14.11 percentage points over Qwen3-30B-A3B-Thinking-2507, compared with +9.94 for same-data SFT. Across eight OOD benchmarks, its average change is +0.21 versus −9.83 for SFT, although its SWE-Bench Verified score is lower than SFT: 32.67 versus 37.40.
- At matched SWE-Bench Verified accuracy of 32.67%, PivotRL uses approximately 133K rollout turns versus approximately 542K for end-to-end RL, with approximately 5.5× less cumulative rollout time on the same number of compute nodes. The theory explains idealized statewise updates; the experiments demonstrate OOD retention and rollout efficiency under the reported conditions, not a general guarantee of capability preservation or lower total training cost.
1. Introduction
The paper addresses the cost–generalization trade-off in post-training agents that interact with tools, code repositories, terminals, and search environments. SFT efficiently reuses fixed demonstrations but can degrade capabilities outside the training domain; end-to-end RL uses on-policy interaction but repeatedly generates expensive multi-turn trajectories.
- A direct local-RL conversion of demonstrations has two weaknesses: randomly selected turns often yield identical group rewards, and exact-match rewards reject valid alternative actions.
- PivotRL addresses these weaknesses through offline pivot selection and domain-specific local verifiers, conditioning on expert histories while sampling new assistant completions on-policy.
2. Preliminaries and Motivating Observations
The preliminaries distinguish turn-level behavior cloning, end-to-end GRPO, and local RL initialized from expert trajectory states. PivotRL changes the training-state distribution and reward source without requiring complete on-policy trajectories from the initial environment state.
Turn-level agentic training.
An action is the complete assistant completion at a model-call boundary, not an individual token. Its state contains the full preceding interaction history; SFT maximizes the demonstrated action’s likelihood, whereas end-to-end RL samples trajectories and uses group-normalized rewards in a clipped GRPO objective with reference-policy KL regularization.
- Equation (1) subtracts the group reward mean and divides by the group standard deviation plus a numerical stabilizer.
- Identical rewards produce zero normalized advantages and remove the reward-driven update signal. A separate KL term can still contribute an update.
Local RL from expert trajectories.
Naive local RL samples an intermediate expert state, generates an on-policy action, and rewards that action only if it exactly matches the demonstrated continuation. On τ²-Bench, this baseline scores 57.34, below same-data SFT at 58.44, showing that on-policy sampling alone does not improve accuracy under this restrictive credit assignment.
Motivating observations.
The motivating experiments report that 71% of randomly sampled turns on τ²-Bench and SWE-Bench yield zero group-normalized learning signal because their sampled actions uniformly succeed or fail. Exact matching introduces a second bottleneck by rejecting tool calls, shell commands, and search steps that are locally acceptable but textually different.
- The miss rate is the probability that strict matching rejects an action conditional on the functional verifier accepting it; the paper does not report a numerical estimate.
- These observations motivate allocating rollout budget to mixed-outcome turns and rewarding more than a single demonstrated string.
3. PivotRL
PivotRL modifies the naive local-RL baseline through offline assistant-turn filtering and verifier-based local rewards. The pipeline reuses expert prefixes, generates short local continuations, and optimizes a GRPO-style objective.
3.1. Method
Offline turn selection. Every extracted assistant turn becomes a pivot candidate. A frozen reference policy generates K local rollouts per candidate; their verifier rewards estimate the reward mean and variance, and the retained set satisfies σ̂²(s) > 0 and μ̂(s) < λdiff.
- The positive-variance condition removes empirically uniform-success and uniform-failure turns.
- The low-mean condition emphasizes difficult but partially solvable states. Selection occurs before RL, and the retained set is not dynamically refreshed during training.
- The PDF specifies λdiff symbolically but does not give its numerical value or a complete profiling-budget specification.
Verifier-based local reward. The verifier defines a set of locally acceptable actions M(s), with reward rfunc(s, a) = 1[a ∈ M(s)]. Acceptance can use normalized string/schema checks, task-specific equivalence rules, or a lightweight LLM judge rather than exact demonstration matching.
- Local acceptance need not establish semantic equivalence or full task success. In SWE-Bench training, the verifier checks only the tool-call name, not its arguments or the resulting patch.
- Terminal-command verification combines output-schema validation, normalized string similarity, and equivalence-based LLM-as-judge scoring over the command and its immediate effect.
Algorithm 1 samples minibatches from retained pivots, generates groups of local on-policy actions, executes the short rollout needed for scoring, and computes group-normalized advantages. Equation (7) applies clipped importance-weighted GRPO updates with KL regularization toward the reference policy.
- The importance ratio compares the current policy with the policy that generated each sampled action.
- Conditioning on stored expert histories avoids regenerating the full prefix online, but demonstration collection and offline pivot profiling remain prerequisites.
3.2. Theoretical analysis for PivotRL
Proposition 3.1 establishes that identical binary rewards yield zero group-normalized advantages. Theorem 3.2 relates reward variance to an idealized statewise update: along the exponentially tilted KL-regularized policy path, the population GRPO score equals the reward standard deviation divided by β².
- The Fisher norm of the statewise reward objective’s natural gradient equals the reward standard deviation.
- At fixed β, larger reward variance therefore gives a larger natural-gradient norm and population GRPO score along this path.
- The score is defined by reward normalization when variance is positive; the proof treats the zero-variance case separately, where the natural gradient is zero. This is not an analysis of complete finite-sample, clipped optimization in a parameter-shared language model.
Theorem 3.3 analyzes binary functional rewards under a fixed state distribution and KL regularization with β > 0. If the reference policy assigns acceptable actions mass ρ(s), the statewise optimum assigns them mass qβ(s) = ρ(s)e^(1/β) / [(1 − ρ(s)) + ρ(s)e^(1/β)], while preserving reference conditional distributions within the acceptable set and its complement.
- For 0 < ρ(s) < 1, acceptable-action mass strictly increases; if every action is acceptable, the optimum remains the reference policy.
- Among distributions with acceptable-action mass qβ(s), the block-rescaled optimum uniquely minimizes KL divergence to the reference. The population result holds on states covered by the fixed distribution, d-almost surely.
- Probability ratios are preserved within each block, not across acceptable and unacceptable actions.
Interpretation. The analysis assumes finite action spaces, nonempty acceptable sets, and positive reference probability for every action. The proposed connection to OOD retention additionally assumes that actions outside the acceptable set are task-unrelated; preserving their within-state ordering does not prove preservation on unseen states or under shared neural-network parameter updates.
4. Experiments
Setup. Four models are trained separately for conversational tool use, software engineering, terminal control, and web browsing, then evaluated on τ²-Bench, SWE-Bench Verified, Terminal-Bench, and BrowseComp respectively. All start from Qwen3-30B-A3B-Thinking-2507 and use Nemo-RL for optimization and Nemo-Gym for environment rollouts; each SFT–PivotRL comparison shares the base model, prompts, and expert trajectories.
- OOD evaluation covers IFBench, AIME25, MATH500, LiveCodeBench, Scicode, MMLU-Pro, MMLU-ProX, and WMT24++.
- The single-domain comparisons control demonstration data but are not equal-total-compute comparisons against SFT.
- The paper states that performance scores and differences are reported in percentage points.
4.1. In-Domain and OOD Accuracy
Table 1 shows PivotRL outperforming same-data SFT on three of four in-domain benchmarks: τ²-Bench reaches 63.81 versus 58.44, Terminal-Bench 20.00 versus 13.75, and BrowseComp 11.30 versus 1.50. SWE-Bench Verified is the exception, at 32.67 versus 37.40, so the +4.17 average advantage over SFT is not a uniform domain-level gain.
Tables 2 and 3 quantify capability retention across the evaluated OOD benchmarks. Averaged across four training runs and eight benchmarks, PivotRL changes by +0.21 relative to Base, whereas SFT changes by −9.83; the largest individual PivotRL regression is −3.12 on AIME25 after terminal-domain training.
- Terminal-domain SFT produces the largest regressions: AIME25 falls from 86.04 to 21.56, MATH500 from 98.05 to 63.55, and WMT24++ from 36.97 to 6.31.
- The corresponding PivotRL scores are 82.92, 98.00, and 36.48, showing that the aggregate retention result also holds in this particularly adverse training domain.
- The +10.04 OOD advantage over SFT follows from Table 2 and agrees with the conclusion. The introduction’s −9.48 SFT change is inconsistent with the tabulated −9.83.
Averaging OOD changes across the four training runs and eight benchmarks gives −9.83 for SFT and +0.21 for PivotRL. PivotRL’s small positive and negative changes indicate retention near Base rather than improvement on every benchmark.
The per-domain breakdown identifies terminal-domain SFT as the source of the largest OOD regressions, including a 64.48-point AIME25 drop. PivotRL avoids comparable regressions in all four training domains, while still showing small individual score decreases.
4.2. Comparison to End-to-End RL on SWE-Bench
Figure 1 compares rollout efficiency on SWE-Bench using the same base model, OpenHands evaluation harness, and number of compute nodes. At 32.67% accuracy, PivotRL reaches step 130 with approximately 133K rollout turns, while end-to-end RL reaches approximately step 72 with approximately 542K turns: approximately 4× fewer turns and approximately 5.5× less cumulative rollout time.
- PivotRL uses batch size 1024, comprising 64 prompts × 16 generations, with one turn per sample.
- End-to-end RL uses batch size 512, comprising 16 prompts × 32 generations, with 12–25 turns per trajectory.
- This is a matched-accuracy rollout comparison. End-to-end RL continues to higher accuracy in Figure 1, and the reported savings do not include a complete accounting of demonstration generation and offline profiling.
At matched accuracy near 32.67%, PivotRL reaches the target with about 4× fewer rollout turns and 5.5× less cumulative rollout time on equal compute-node counts. End-to-end RL continues to higher accuracy, so the figure establishes efficiency at a shared target rather than higher final accuracy.
4.3. Ablation Study
Table 4 reports 63.81 on τ²-Bench for full PivotRL and 59.68 when filtering is removed but functional rewards are retained. The unfiltered strict-reward configuration scores 57.34, below same-data SFT at 58.44.
- The functional-reward comparison isolates the benefit of pivot filtering.
- Although the paper describes removing components one at a time, the strict-reward row uses unfiltered Dcand rather than filtered Dadv. Comparing 63.81 with 57.34 therefore changes both selection and reward; comparing 59.68 with 57.34 more directly tests functional versus strict rewards without filtering.
Filtered functional-reward training scores 63.81, versus 59.68 without filtering and 57.34 with unfiltered strict rewards. The strict-reward row also removes filtering, so its comparison with full PivotRL does not isolate reward design alone.
Figures 2 and 3 show higher validation accuracy and higher per-batch reward standard deviation with filtered pivots than with random candidate sampling. Appendix Table 6 reports best-checkpoint average scores of 59.68 for random candidates and 63.81 for low-reward-mean pivots, compared with 58.44 for SFT.
- The gain is not uniform across τ²-Bench subdomains: filtered pivots score 63.74 on Telecom, below SFT at 65.79, while improving Airline and Retail.
- Random sampling’s reward standard deviation declines but remains visibly nonzero; the plot supports more informative sampling with filtering, not a complete collapse of the random-sampling signal or proof that variance alone causes the accuracy gain.
After its initial rise, the filtered-pivot curve generally exceeds random-candidate training and the SFT reference. Checkpoint fluctuations show that the reported best-checkpoint score is not maintained throughout training.
Filtered pivots maintain reward standard deviation close to 0.50, whereas random sampling produces a lower, declining value that remains nonzero. This supports the claim that filtering sustains more mixed-outcome sampling, without showing that random sampling loses all learning signal.
Filtered pivots achieve the highest listed τ²-Bench average at 63.81, with Airline and Retail improvements over both SFT and random candidates. Telecom remains below SFT, so the average benefit is not uniform across subdomains.
4.4. Integration into Large-Scale Post-Training
PivotRL is used for agentic environments in the Nemotron-3-Super post-training pipeline, alongside other RL environments for reasoning and chat. Table 5 reports changes from the SFT checkpoint to the subsequent mixed RL stage: τ²-Bench 48.00 to 64.00, SWE-Bench Verified 12.87 to 61.33, Terminal-Bench 1.1 Core 23.33 to 34.17, and BrowseComp 13.03 to 25.04.
- These results document integration into a larger post-training recipe.
- Because the stage includes other RL environments, the table does not isolate PivotRL’s contribution to the full checkpoint improvement.
All four agentic scores increase after the Nemotron-3-Super RL stage, including SWE-Bench Verified from 12.87 to 61.33. Since the stage combines PivotRL agentic environments with reasoning and chat environments, the results document pipeline integration rather than an isolated PivotRL effect.
5. Related Works
The related-work discussion connects agentic policy optimization and verifiable rewards with methods that reuse offline demonstrations for online learning. PivotRL combines selected assistant-turn states with local verifier-derived credit.
5.1. Agentic LLMs and Training Recipes
Agentic LLMs interleave language with environment-grounded actions, while RL supports exploration and multi-turn credit assignment. PivotRL draws on RL with verifiable rewards and on-policy learning, converting supervised trajectories into turn-level reward-bearing states without training a separate reward model.
- Verification is not exclusively programmatic: the method permits LLM judges, and the Terminal-Bench implementation uses one.
5.2. Adapting Behavior Cloning into RL
The paper relates its approach to catastrophic forgetting in SFT, horizon-dependent imitation-learning errors, and demonstration-assisted RL. Prior methods mix expert demonstrations with RL updates, reuse partial prefixes, or convert SFT tokens into rollout rewards; PivotRL applies this SFT-to-RL bridge to multi-turn tool use through selected assistant states and functional rewards.
6. Conclusion
The conclusion attributes PivotRL’s gains to broader action coverage through on-policy rollouts and verifier-derived rewards. The evaluations support an average +4.17 in-domain advantage and +10.04 OOD-retention advantage over same-data SFT, plus lower rollout cost than end-to-end RL at matched SWE-Bench accuracy.
- Proposed extensions include process reward models and online reward profiling such as dynamic sampling.
- The conclusion also proposes non-programmatic verification, although LLM-as-judge scoring is already described for Terminal-Bench in Appendix A.2.
Appendix
- A. Experiment Details and A.1. Terminology Conventions distinguish extracted pivot candidates from the offline-filtered pivots used in training. Actions are complete assistant completions at model-call boundaries, including natural-language-only turns when applicable; OOD change is measured relative to Base on benchmarks outside the training domain. The random set uses all candidates, while the low-reward-mean set requires mixed outcomes and a gap between demonstrated-action reward and reference-policy mean.
- A.2. Domain-Specific Training Details — τ²-Bench. Candidates are extracted, profiled, and filtered before RL in all four domains. The conversational tool-use dataset contains 281,774 trajectories across 838 domains, generated with a synthetic pipeline similar to ToolAce, Kimi-K2, and DeepSeek-V3.2; assistant turns can contain language, tool calls, or both. The environment and data are released through Nemo-Gym and Nemotron-Post-Training-v3.
- Terminal-Bench. Resolved trajectories come from Qwen/Qwen3-Coder-480B-A35B-Instruct and moonshotai/Kimi-K2-Instruct using Terminus-1 and Terminus-2. Bash-command candidates are filtered and deduplicated to approximately 20,000 samples, with local interchangeability assessed through schema validation, normalized string similarity, and equivalence-based LLM scoring.
- SWE-Bench Verified. Internal trajectories are generated with OpenHands, OpenCode, and Codex on SWE-Gym and R2E-Gym tasks using MiniMaxAI/MiniMax-M2.5; non-error tool-call candidates are filtered to Dadv, yielding 87,718 samples. Evaluation uses mean@3 with OpenHands, whereas the training verifier matches tool-call names only and does not score arguments or patch quality. Final task success is determined by the full SWE-Bench evaluation harness.
- BrowseComp. Browsing trajectories are generated from a multi-hop question-answer dataset using an online search engine and DeepSeek-V3.2, yielding 13,215 samples. Actions include issuing queries, opening results, and gathering evidence; the appendix does not specify the domain’s local verifier in comparable detail.
- A.3. Detailed Ablation Analyses and A.3.1. Effect of Turn Selection Strategy on τ²-Bench compare unfiltered random candidates with mixed-outcome, low-reward-mean pivots defined by Equation (5). Table 6 reports the best checkpoint over training: average accuracy is 58.44 for SFT, 59.68 for random candidates, and 63.81 for filtered pivots.
- B. Omitted Proofs — B.1. Proof of Theorem 3.2 derives the natural gradient as centered reward, equates its squared Fisher norm with reward variance, and differentiates the exponentially tilted KL path to obtain the population GRPO score. B.2. Proof of Theorem 3.3 decomposes KL into conditional-distribution terms and a Bernoulli mass term; strict convexity determines the optimal acceptable-action mass, and blockwise rescaling preserves within-block probability ratios. The proof also treats the boundary cases where no actions or all actions are acceptable.
Brief Thoughts
The main contribution is the same-data evidence that selected local on-policy training can acquire agentic skills while retaining capabilities that SFT degrades in these experiments. The matched-accuracy SWE-Bench comparison provides a concrete rollout-efficiency result, but PivotRL’s lower score than SFT in that domain shows that retention and efficiency do not necessarily maximize in-domain accuracy.
The central limitations are coarse local credit, fixed expert-state coverage, and incomplete reproducibility and cost reporting. Tool-name-only coding rewards do not establish semantic action equivalence, and the statewise theory does not guarantee neural-model OOD retention. The PDF does not report repeated-run uncertainty, complete optimization hyperparameters, or an end-to-end cost accounting that includes offline profiling.