Experiential Reinforcement Learning | Summary
15 Feb 2026 | Paper Review Reinforcement Learning Self-Reflection Self-DistillationContents
- Summary
- 1 Introduction
- 2 Experiential Reinforcement Learning (ERL)
- 3 Experiment
- 4 Result and Discussion
- 5 Related Work
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Experiential Reinforcement Learning.
- 2026-02-15 (arXiv)
- Shi, Taiwei, Chen, Sihao, Jiang, Bowen, Song, Linxin, Yang, Longqi, Zhao, Jieyu.
- University of Southern California, Microsoft, University of Pennsylvania
- Paper
Summary
- Experiential Reinforcement Learning (ERL) trains language-model agents through an experience–reflection–consolidation loop. After a suboptimal initial attempt, the model generates feedback-conditioned guidance, retries the task, and internalizes positive-reward retries through selective self-distillation.
- ERL achieves higher final evaluation reward than GRPO-based RLVR in all six combinations of three tasks and two backbones. The largest absolute gain is on Qwen3-4B-Instruct-2507 in Sokoban, from 0.06 to 0.87; HotpotQA gains are smaller, from 0.45 to 0.56 for Qwen and from 0.47 to 0.50 for Olmo.
- Reflection-conditioned retries generally outperform initial attempts during training, and removing structured reflection lowers final reward in every evaluated setting. Memory has mixed effects: it improves some results, leaves others unchanged, and lowers Olmo’s Sokoban reward relative to the no-memory variant.
- Deployment does not require the training-time reflection pass or feedback from a previous attempt. The evidence covers two backbones and three tasks, including deterministic 8-action control episodes, without reported seed-level uncertainty; the PDF also contains unresolved rollout-allocation and memory-threshold inconsistencies.
1 Introduction
Sparse terminal rewards identify unsuccessful trajectories without specifying how an agent should revise its decisions. ERL makes corrective reasoning an explicit training output rather than leaving behavioral revision entirely to scalar-reward optimization.
- The motivation draws on experiential learning: observing outcomes, reflecting on failures, revising behavior, and consolidating successful corrections.
- Figures 1 and 2 distinguish imitation, reward-driven trial-and-error, and reflection-guided learning. They are conceptual illustrations, not evidence that RLVR necessarily forgets failed attempts.
- The stated contributions are an explicit training loop, internalization into a reflection-free policy, and empirical improvements across control and tool-using reasoning tasks.
The three panels distinguish imitation of examples, scalar-reward optimization, and reflection followed by internalization. ERL's added path transfers reflection-conditioned behavior to the original-context policy; the schematic illustrates the intended mechanism rather than measuring its effectiveness.
The ERL row follows a Sokoban-like failure with hypotheses about walls, controllable entities, and pushing behavior, then depicts consolidation across attempts. The repeated failures in the RLVR row are a motivational schematic, not measured trajectories from the single-box experiments.
2 Experiential Reinforcement Learning (ERL)
Given task x, the policy samples y^(1) from πθ(· | x) and receives textual feedback f^(1) and reward r^(1). When the reflection gate activates, the same model generates Δ conditioned on the task, initial trajectory, feedback, reward, and memory m, then samples y^(2) from πθ(· | x, Δ).
- The reflection receives the second attempt’s reward, r̃ = r^(2), linking its optimization to downstream task performance.
- Cross-episode memory supplies prior corrective guidance during reflection generation; the implementation uses plain-text system-prompt guidance rather than retrieval-based memory.
- Figure 3 separates reward-driven learning of attempts and reflections from supervised internalization. Algorithm 1 is simplified; Appendix A provides the gated experimental version.
The upper path uses reinforcement learning for the initial attempt, reflection, and revised attempt with the same policy. The lower path learns the revised output from the original input without reflection, separating training-time contextual guidance from parameter-level internalization.
The policy objective is Lpolicy(θ) = −E[A log πθ(y | x, ·)], where y can be an initial attempt, reflection, or revised attempt and A is estimated from its associated reward. Internalization uses Ldistill(θ) = −E[I(r^(2) > 0) log πθ(y^(2) | x)], removing reflection from the conditioning input.
- For binary-reward control tasks, the distillation filter selects successful retries. For HotpotQA, it also admits partially correct retries with positive reward.
- The filter does not require the second attempt to improve on the first; its stated condition is positive second-attempt reward.
- The deployment policy is trained to reproduce revised behavior from the original task context without first generating a reflection.
2.1 Comparison to Standard RLVR
Standard RLVR optimizes outcome rewards without an explicit feedback-conditioned reflection-and-retry stage. ERL retains reward-driven policy optimization but adds reflection and retry to the training trajectory, then distills positive-reward retries into the original-context policy.
- Environmental rewards still supervise the attempts and determine the reflection reward; verbal reflection does not replace verifiable task evaluation.
- Memory reuses guidance in the training context, whereas internalization updates model parameters so deployment does not depend on that guidance.
3 Experiment
The experiments compare ERL with standard RLVR on Frozen Lake, Sokoban, and HotpotQA. Training two backbones under both paradigms produces six task–model comparisons.
3.1 Task
Frozen Lake and Sokoban provide terminal reward +1 for completing the objective and 0 otherwise. Explicit game rules and symbol definitions are withheld, so the agent must infer dynamics from grid observations, available actions, and textual transition feedback.
- Both control tasks use deterministic transitions and an 8-action budget. Sokoban instances contain one box and one goal and are accepted only when BFS finds a solution of at most 8 moves.
- Each control task uses 10,000 procedurally generated training instances and 100 disjoint evaluation instances generated by the same process.
- The initial prompt nevertheless supplies a structured reasoning procedure: state analysis, strategic-value assessment, and consequence predictions for two candidate actions. Withholding game rules therefore does not mean withholding all reasoning guidance.
HotpotQA is adapted into tool-assisted multi-hop question answering following Search-R1. The agent retrieves Wikipedia evidence and places its final answer inside \boxed{}; episodes permit up to 5 interaction turns.
- Exact matches receive reward 1.0. Otherwise, reward equals token-level F1 when F1 ≥ 0.3 and is 0 below that threshold.
- This scoring gives partial credit for some incorrect answers, unlike the binary control rewards.
- The environment configuration does not report HotpotQA training or evaluation sample counts.
3.2 Models and Baselines
The backbones are Olmo-3-7B-Instruct and Qwen3-4B-Instruct-2507, with GRPO used for both ERL and RLVR. Clipping, KL regularization, and importance sampling stabilize optimization; the authors also apply these techniques during internalization because its targets are off-policy relative to the original-context policy.
- Section 3.2 assigns RLVR 10 rollouts per task and ERL half as many per attempt, implying 5. Appendix C instead specifies 4 samples per prompt for each ERL attempt.
- The allocation is intended to match compute despite the reflection step, but the PDF provides no token-level compute accounting that resolves the discrepancy.
- All runs use the rLLM training stack on a single node with 8 H100s.
4 Result and Discussion
The results distinguish final evaluation reward, validation reward against wall-clock time, and training reward before versus after reflection. Table 1 and Figures 4–6 address final policy performance, observed training speed, and immediate retry improvement, respectively.
- Curves use a trailing moving average over 5 points unless otherwise noted.
- Figure 4 shows an overall wall-clock advantage, especially in the control tasks, but ERL is not ahead at every point.
- The reported table and figures provide no confidence intervals or variability across independent training seeds.
Elapsed training hours expose practical learning-speed differences rather than differences in update count alone. Control-task panels show large late-stage separations, particularly Qwen on Sokoban; HotpotQA margins are smaller, and Olmo's RLVR curve initially leads before ERL overtakes it.
4.1 Performance Across Tasks
For Qwen3-4B-Instruct-2507, RLVR → ERL rewards are 0.86 → 0.94 on FrozenLake, 0.45 → 0.56 on HotpotQA, and 0.06 → 0.87 on Sokoban. For Olmo-3-7B-Instruct, the corresponding rewards are 0.39 → 0.66, 0.47 → 0.50, and 0.04 → 0.20.
- The introduction’s gains of up to +81%, +27%, and +11% denote absolute reward differences of 0.81, 0.27, and 0.11, not relative percentage increases.
- The authors attribute larger control-task gains to inferring unknown dynamics and correcting compounding errors. Task difficulty and reward density are not independently manipulated, so this explanation remains an interpretation.
- Although the discussion emphasizes long-horizon planning, the configured control episodes permit only 8 actions. Improvements on two backbones also do not establish architecture-independent benefits.
ERL has higher final reward in all six task–model comparisons. Effect sizes vary considerably: Qwen's Sokoban gain is 0.81, whereas Olmo's HotpotQA gain is 0.03; the bars contain no uncertainty intervals.
4.2 Learning Efficiency and Optimization Dynamics
Figure 4 shows faster improvement and higher late-stage validation reward for ERL overall. The largest separation appears in Sokoban, where Qwen’s ERL reward rises sharply while RLVR remains near zero.
- The authors interpret the curves as feedback-conditioned revision directing exploration toward more useful trajectories rather than relying only on terminal-reward propagation.
- HotpotQA has smaller margins; Olmo’s RLVR curve is initially ahead before ERL overtakes it.
- Wall-clock curves document observed training-speed differences but do not isolate reflection, distillation, memory, or rollout allocation as separate causes.
4.3 Mechanistic Role of Reflection
Figure 6 shows that post-reflection training reward generally exceeds pre-reflection reward across both backbones and all three tasks. The sustained separation is consistent with feedback-conditioned guidance improving within-episode retries.
- Pre-reflection reward also improves during training, indicating better initial-attempt behavior rather than gains confined to retries.
- Retries receive information unavailable to initial attempts, so this comparison alone does not distinguish structured reflection from generic contextual reuse. The no-reflection ablation provides the more relevant control.
Post-reflection training reward generally exceeds pre-reflection reward, consistent with immediate benefits from feedback-conditioned retry. Rising pre-reflection curves also indicate better initial-attempt behavior, but the plot does not isolate internalization from the other ERL components.
4.4 Ablation Study: Memory and Reflection Mechanisms
The no-memory ablation retains reflection and retry but disables cross-episode reuse. The no-reflection ablation retains two attempts and supplies the full initial trajectory with a generic improvement instruction, documented in Table 9, instead of synthesized reflective guidance.
- For Qwen, removing reflection lowers reward from 0.94 to 0.60 on FrozenLake, 0.56 to 0.48 on HotpotQA, and 0.87 to 0.59 on Sokoban.
- For Olmo, removing reflection lowers the corresponding rewards from 0.66 to 0.54, 0.50 to 0.46, and 0.20 to 0.06.
- Figure 7 shows slower improvement without memory and a substantially weaker curve without reflection for Qwen on FrozenLake.
Memory is less consistently beneficial than structured reflection. Removing it changes Qwen’s rewards to 0.86, 0.56, and 0.87, and Olmo’s rewards to 0.64, 0.47, and 0.24 on FrozenLake, HotpotQA, and Sokoban, respectively.
- Olmo’s Sokoban reward rises from 0.20 to 0.24 without memory, ruling out a universal final-performance advantage from memory reuse in these results.
- The authors suggest that inaccurate early reflections can persist and impede recovery, but do not measure reflection accuracy or error propagation. Their reference to stochastic environments does not describe the configured deterministic control tasks.
- No reported ablation removes internalization alone, so the deployment-policy gains cannot be attributed specifically to selective distillation.
Removing structured reflection lowers reward in every setting, whereas removing memory has smaller and mixed effects. Olmo's Sokoban result is 0.24 without memory versus 0.20 with full ERL, showing that memory does not universally improve final performance.
For Qwen on FrozenLake, full ERL improves faster and reaches a higher validation-reward plateau than either ablation. Removing memory preserves much of the improvement but delays progress; generic trajectory-conditioned retry produces a weaker curve and finishes below RLVR in this panel.
The retry receives observations, actions, rewards, and feedback from the initial attempt but no synthesized reflection. This controls for reuse of prior experience more directly than an independent second sample, while changing how that experience is represented relative to full ERL.
5 Related Work
The related-work discussion connects ERL to RLHF, verifiable-reward reasoning, tool-using agents, hindsight experience replay, and learning from self-generated interaction traces. The closest methodological comparisons are inference-time reflection and feedback-conditioned teacher–student distillation.
- Inference-time methods such as Self-Refine and Reflexion typically retain critique or memory at deployment, whereas ERL aims to absorb revised behavior into model parameters.
- Concurrent work on self-distillation and textual feedback also transfers feedback-conditioned behavior to a deployment policy. ERL emphasizes reward-trained reflection within an attempt–retry trajectory, selective internalization, and cross-episode memory.
6 Conclusion
The conclusion presents explicit reflection and selective internalization as a way to convert environmental feedback into deployable behavioral changes. The experiments show higher final reward and an overall wall-clock learning advantage in the evaluated settings without requiring reflection at inference.
- Continual adaptation remains a proposed direction: the experiments do not test lifelong learning, retrieval-based memory, or transfer to substantially different environments.
Appendix
- A Full Algorithm and Gated Reflection: Algorithm 2 updates the initial attempt and triggers reflection only when r^(1) < τ, with τ = 1. The authors report qualitatively that reflecting on successful attempts encouraged instance-specific reward hacking and that reflection-conditioned trajectories could dominate early optimization; they do not provide a quantified gating ablation.
- The memory-storage condition in Algorithms 1 and 2 is r^(2) > τ. With τ = 1 and documented rewards bounded above by 1.0, this strict inequality cannot be satisfied, leaving an unresolved inconsistency between the pseudocode and the reported memory experiments.
- Appendix A describes memory as plain-text system-prompt guidance updated by direct overwrite, not an accumulating retrieval database. Retrieval-based updates and on-policy distillation are proposed extensions rather than evaluated components; the latter is called reverse KL in the prose, while the displayed KL places the reflection-conditioned policy before the deployment policy.
- B Envrionment Configuration Details — B.1 Frozen Lake: grid size n is sampled uniformly from [2, 9], and frozen-tile probability p from [0.6, 0.85). Start and goal positions are distinct, generation ensures a valid path, and transitions are deterministic; the encoding is A = agent position, B = goal tile, C = hole, and D = safe frozen tile.
- B.2 Sokoban: n is sampled uniformly from [6, 8], with border walls, one box, one goal, and non-overlapping interior placements. BFS accepts layouts with solutions of at most 8 moves; actions are {Up, Down, Left, Right}, invalid moves leave the state unchanged, and the step budget is 8.
- B.3 HotpotQA: retrieval uses the PeterJinGo/wiki-18-corpus Wikipedia corpus, PeterJinGo/wiki-18-e5-index dense indices, intfloat/e5-base-v2 embeddings, and FAISS with multi-GPU support. The tool is local_search with query and top_k parameters; extracted \boxed{} answers are normalized by lowercasing and whitespace canonicalization before scoring.
- C Training Configuration Details: training uses rLLM, GRPO, vLLM, FlashAttention, hybrid engine training, gradient checkpointing, and remove-padding on 8 H100s. The learning rate is 1e-6; training and actor mini-batch sizes are 64; maximum prompt and response lengths are each 8,196 tokens; and the actor maximum token length per GPU is 24,000. Actor parameter/optimizer offload and reference-model parameter offload are enabled.
- Training rollouts use temperature 0.7, GPU memory utilization 0.85, asynchronous generation, and tensor model parallel size 1. Validation uses 4 samples per prompt, temperature 0.7, top-p 0.8, and top-k 20; KL loss and fixed KL control coefficients are 0.001, the actor clipping ratio upper bound is 0.28, and the entropy coefficient is 0.
- Appendix C specifies 10 samples per prompt for RLVR and 4 per attempt for ERL, evaluation every 5 iterations, and manual early stopping upon convergence. Rejection sampling and stepwise advantage estimation are disabled. Tables 2–9 supply initial prompts, reflection prompts, example tasks, and the generic no-reflection retry template.
Brief Thoughts
The no-reflection ablation provides the clearest evidence for the value of explicit reflective guidance: it outperforms supplying the previous trajectory with a generic retry instruction. The deployment-policy results matter separately because they do not require the training-time reflection loop at inference.
The experimental scope limits broader conclusions. Two backbones, 100 evaluation layouts per control task, deterministic 8-action episodes, and no seed-level uncertainty do not establish robustness across architectures, long horizons, or out-of-distribution dynamics.
Reproducibility requires resolving the strict memory threshold and the 5-versus-4 rollout discrepancy. Compute accounting, independent-run variability, and an ablation that removes only internalization would make component-level attribution more precise.