Online Experiential Learning for Language Models | Summary
17 Mar 2026 | Paper Review Online Learning On-Policy Distillation Agent LearningContents
- Summary
- 1 Introduction
- 2 Preliminary: Online Learning
- 3 Online Experiential Learning
- 4 Experiments
- 5 Related Work
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Online Experiential Learning for Language Models.
- 2026-03-17 (arXiv)
- Ye, Tianzhu, Dong, Li, Dong, Qingxiu, Wu, Xun, Huang, Shaohan, Wei, Furu.
- Microsoft Research
- Paper
- Project Page
Summary
- Online Experiential Learning (OEL) turns deployment trajectories into model updates without reward supervision. It extracts transferable experiential knowledge from textual interaction histories and internalizes that knowledge through on-policy context distillation, without server-side access to the interaction environment.
- Experiments use Frozen Lake and Sokoban in TextArena with Qwen3 thinking models of 1.7B, 4B, and 8B parameters and Qwen3-4B-Instruct-2507. Successive rounds improve game pass rates; for Qwen3-1.7B on Frozen Lake, average per-turn response length falls to roughly 70% of its initial value by the third iteration.
The paired plots show rising pass rate and falling normalized response length across alternating extraction and consolidation phases. Consolidation arrows indicate changes in the context-free model after parameter updates, while extraction curves show the effect of accumulating knowledge in context.
- Ablations favor extracted knowledge over raw trajectories and self-derived knowledge over knowledge from a larger model. On-policy consolidation also retains IF-Eval accuracy better than an off-policy alternative, although the evaluation covers only two small game environments and one out-of-distribution benchmark.
1 Introduction
The paper addresses learning from deployment experience when the training server cannot access user-side environments and interactions provide textual feedback rather than scalar rewards. OEL uses these histories to construct training signals instead of requiring human annotations or a reward function for each environment.
- The user side collects interaction trajectories; knowledge extraction and parameter updates occur on the server.
- The procedure requires no human annotations, reward model, or verifiable reward function. Game success is an evaluation metric, not a training reward.
- The paper is Part II of the Experiential Learning series, following On-Policy Context Distillation for Language Models. Code is listed at aka.ms/oel-code.
2 Preliminary: Online Learning
The paper frames online learning as repeated deployment, experience collection, and server-side updating, rather than training only on a preconstructed dataset. Figure 2 separates user-side interaction from server-side learning; its open-world framing describes the proposed deployment paradigm, not the setting directly tested in the experiments.
- The lack of server-side environment access does not eliminate environment interaction: the deployed model still needs to collect trajectories.
- The experiments instantiate this architecture with controlled TextArena games rather than real user deployment.
The diagram separates user-side interaction from server-side updating with collected experience. Its open-world labels describe the proposed deployment paradigm; they should not be read as evidence that the experiments use real-world environments rather than controlled games.
3 Online Experiential Learning
OEL separates experience processing into knowledge extraction and weight consolidation. Figure 3 shows how collected trajectories supply experiential knowledge for a teacher and partial interaction histories from which the student generates fresh single-turn responses.
- The teacher receives experiential knowledge in context; the student receives only the interaction prefix.
- Reverse KL divergence supplies token-level supervision on student-generated responses, so consolidation requires no further environment interaction.
3.1 Extract Experiential Knowledge from User Trajectories
The deployed model collects trajectories alternating textual environment feedback and model actions. By default, the same model extracts new knowledge from each trajectory while conditioning on accumulated knowledge: e′ᵢ ∼ πextract(· | τᵢ, eᵢ₋₁), eᵢ = [eᵢ₋₁; e′ᵢ], with e₀ = ∅.
- Accumulation concatenates existing knowledge and newly extracted content; extraction uses no ground-truth labels.
- The prompts request transferable rules and strategies rather than descriptions specific to one board.
- For Qwen3-1.7B, previously accumulated knowledge is omitted from the extraction context because the authors found that the smaller model struggled to use long contextual information.
3.2 Consolidate Experiential Knowledge into Model Weights
The experiments use the Cross-Trajectory variant: one trajectory set supplies accumulated knowledge, while a separate set supplies consolidation examples. Each training example is a partial rollout prefix ending at an environment observation, immediately before the next model response.
- Repeating extraction with different random seeds produces a collection of knowledge contexts.
- Each training step samples a prefix and a knowledge context separately, allowing knowledge to be applied beyond its originating trajectory.
- The server generates a new single-turn response without executing it in the environment.
The student samples a response without experiential knowledge and minimizes average token-level reverse KL divergence to a knowledge-conditioned teacher: L(θ) = Eₓ∼D,e∼C,y∼πθ(·|x)[|y|⁻¹ Σₜ₌₁^{|y|} DKL(πθ(· | x, y<t) ∥ πteacher(· | e, x, y<t))]. The teacher is a frozen copy of the model at the start of the consolidation round.
- Both models are evaluated on the same student-generated response prefixes, but only the teacher receives the knowledge context.
- Training is on-policy with respect to newly generated response tokens; the interaction histories are precollected.
- The implementation approximates vocabulary-level reverse KL using the top 256 tokens ranked by student probability.
Collected histories supply extracted knowledge and partial rollout prefixes. The student generates from a prefix alone, while the teacher receives knowledge and the prefix; reverse KL transfers the teacher’s contextual information without requiring the server to execute new actions in the environment.
3.3 Online Learning Process
After consolidation, the updated model is redeployed to collect fresh extraction and training trajectories. Algorithm 1 repeats collection, extraction, consolidation, and redeployment, setting the extractor and frozen teacher to the current model in each round.
- The authors propose that improved behavior produces more informative experience for subsequent updates.
- Consolidation removes the need to retain extracted knowledge in the deployed model’s inference context.
4 Experiments
The experiments examine iterative improvement, response-token efficiency, out-of-distribution instruction-following retention, model-size effects, and the source and representation of experiential knowledge. Game pass rate is the main task metric; IF-Eval measures retention in the on-policy versus off-policy comparison.
4.1 Setup
Datasets and Models: Frozen Lake uses a 3 × 3 grid with two holes, and Sokoban uses a 6 × 6 grid with one box. The models are Qwen3-1.7B, Qwen3-4B, and Qwen3-8B in thinking mode, plus the non-thinking Qwen3-4B-Instruct-2507.
- Initial prompts replace explicit game mechanics with a general objective, the player identifier, movement actions, and required action syntax.
- TextArena still returns textual action outcomes and updated maps, allowing models to infer mechanics from feedback.
- Evaluation uses a held-out split of 128 game maps, with pass rates averaged over 10 random seeds.
Extraction Stage: Structured knowledge retains entries prefixed with “– EXPERIENCE ITEM:”, accumulates over n = 25 or n = 50 trajectories, and has a maximum length of 8192 tokens. Unstructured knowledge uses n = 15 and a maximum length of 2048 tokens; both formats repeat accumulation K = 10 times with different random seeds.
- Accumulated content exceeding the applicable length limit is truncated.
- For thinking extractors, the answer portion is retained as knowledge and the reasoning portion is removed.
- Knowledge is taken at a fixed accumulation step across rounds rather than selected by reward or best observed performance.
Consolidation Stage: Each round uses 20 or 100 training steps with 64 game samples per step, corresponding to 1280 or 6400 trajectory samples. Collected interactions span up to 5 turns with a maximum response length of 1024 tokens per turn; evaluation uses the final-step checkpoint without checkpoint selection.
- Learning rates are selected from {1e−6, 5e−6}, sampling temperature is 0.7, and hyperparameters remain fixed across rounds for each model-task pair.
- Qwen3-1.7B on Frozen Lake uses unstructured knowledge, n = 15, 2048 tokens, learning rate 5e−6, and 20 steps.
- On Frozen Lake, Qwen3-4B and Qwen3-8B use structured knowledge with n = 25 and n = 50 and 20 and 100 training steps, respectively. Qwen3-4B-Instruct-2507 on Sokoban uses structured knowledge, n = 50, and 100 steps; all three configurations use an 8192-token knowledge limit and learning rate 1e−6.
4.2 OEL Enables Online Learning
Figure 4 shows repeated gains on Frozen Lake with Qwen3-1.7B and on Sokoban with Qwen3-4B-Instruct-2507. Knowledge accumulation improves in-context performance but eventually saturates; consolidation improves performance without the knowledge context and provides a stronger starting model for the next round.
- The authors attribute gains beyond the teacher’s measured in-context performance to generalization from consolidation data, using the teacher’s dense token-level supervision.
- Transparent continuation curves show the limits of continued context accumulation without the corresponding consolidation update.
- The evidence supports improvement over the displayed rounds, not indefinite improvement under arbitrary deployment conditions.
Both game panels show stronger context-free starting policies after consolidation and further gains during subsequent extraction. Transparent continuations show that extending context accumulation alone can plateau or deteriorate, whereas alternating it with weight updates supports improvement over the displayed rounds.
4.3 OEL Improves Token Efficiency
For Qwen3-1.7B on Frozen Lake, Figure 5 reports average per-turn response length falling to roughly 70% of its initial value by the third iteration. Shorter responses accompany higher pass rates, and consolidation retains the reduction without requiring experiential knowledge at inference time.
- The efficiency metric measures generated response length, not end-to-end latency or total learning cost.
- The normalized response-length metric excludes the additional computation required for extraction and consolidation.
Average per-turn response length for Qwen3-1.7B on Frozen Lake approaches roughly 70% of its initial value by the third iteration. The updated model retains shorter responses without the knowledge context; the plot measures generated-token length rather than latency or total learning cost.
4.4 OEL Mitigates Catastrophic Forgetting
Figure 6 compares on-policy and off-policy context distillation for Qwen3-1.7B on Frozen Lake. On-policy training reaches higher game pass rates and keeps IF-Eval prompt-level strict accuracy closer to the initial model, while the off-policy alternative shows a larger decline.
- The off-policy baseline generates responses with the knowledge-conditioned teacher and trains the context-free student using forward KL divergence.
- The comparison concatenates Round 1 and Round 2 consolidation stages, each with 20 gradient steps and batch size 64; saturated portions are omitted and the concatenated curves are smoothed.
- The result demonstrates better IF-Eval retention in this comparison, not preservation of every general language-model capability.
On-policy consolidation reaches a higher game pass rate and shows a smaller IF-Eval decline than off-policy consolidation. The curves concatenate two training stages, omit saturated portions, and apply smoothing, so the horizontal progression is not an unmodified single-run trace.
4.5 Effect of Model Size
Figure 7 reports Frozen Lake performance for Qwen3-1.7B, Qwen3-4B, and Qwen3-8B before training and after two OEL rounds. All three improve, and Round 2 exceeds Round 1 at every scale; Round 2 pass rates increase with model size, while Round 1 results are nonmonotonic.
- The authors attribute stronger later-round performance to higher-quality trajectories and more effective extracted knowledge; the plot does not independently establish this mechanism.
- The comparison is not compute-matched: Qwen3-8B uses 100 consolidation steps, while Qwen3-1.7B and Qwen3-4B use 20, and extraction configurations also differ.
- The broken vertical axis affects the apparent separation between initial and trained performance.
All three model sizes improve after OEL, and their second-round results exceed their first-round results. Second-round performance increases with size, but first-round performance is nonmonotonic; differing extraction settings and training budgets prevent attributing the differences to model capacity alone.
4.6 Analysis
4.6.1 Learning from Experiential Knowledge over Raw Experience: Table 1 evaluates first-round Sokoban performance with Qwen3-4B-Instruct-2507. The no-experience baseline is 7.5%; raw trajectories yield 10.9% in context and 7.8% after consolidation, while extracted knowledge yields 18.2% in context and 21.4% after consolidation.
- Knowledge consolidation improves the baseline by 13.9 percentage points, compared with 0.3 percentage points for raw-trajectory consolidation.
- This ablation supports extraction in the tested configuration; it does not show that all raw-trajectory training methods are ineffective.
Extracted knowledge raises Sokoban pass rate from 7.5% to 18.2% in context and 21.4% after consolidation. Raw trajectories reach 10.9% in context but only 7.8% after consolidation, directly supporting the extraction stage in this configuration.
4.6.2 On-Policy Consistency Between Experiential Knowledge and Policy Model: Table 2 evaluates Qwen3-1.7B on Frozen Lake from a 7.3% baseline. Knowledge from Qwen3-4B yields 18.0% in context and 22.7% after consolidation, while self-derived knowledge yields 23.8% and 31.1%, respectively.
- Self-derived knowledge exceeds larger-model knowledge by 5.8 percentage points in context and 8.4 percentage points after consolidation.
- The authors suggest that larger-model knowledge may encode strategies the smaller model cannot effectively execute.
- The comparison favors self-derived knowledge for this model pairing but does not isolate the cause of the transfer gap.
For Qwen3-1.7B, self-derived knowledge outperforms knowledge from Qwen3-4B both in context and after consolidation. Consolidated pass rates are 31.1% and 22.7%, respectively; the result favors self-derived knowledge in this source-target comparison without identifying the precise cause of the gap.
5 Related Work
The related-work discussion connects OEL to on-policy distillation, context distillation, and learning from interaction experience. OEL combines trajectory-derived knowledge extraction with parameter consolidation in a repeated deployment loop.
On-Policy Distillation
On-policy distillation trains on student-generated responses rather than teacher-generated examples, reducing mismatch between training responses and the student’s inference distribution. OEL uses this mechanism with reverse KL divergence to internalize a knowledge-conditioned teacher’s behavior.
- The paper associates reverse KL with mode-seeking behavior and cites prior work on on-policy distillation and reduced forgetting.
- OEL’s direct retention evidence is the Frozen Lake and IF-Eval comparison.
Context Distillation
Context distillation transfers information supplied in a teacher’s context into student parameters, removing the need for that context at inference time. OEL builds on On-Policy Context Distillation for Language Models rather than introducing context distillation itself.
- The paper contrasts student-generated, reverse-KL training with conventional teacher-generated, forward-KL context distillation.
- OEL derives the guiding context from deployment trajectories and repeats extraction and distillation across rounds.
Learning from Experience
The paper connects OEL to experience-driven agent learning, reflection on failures, and extraction of insights into external memory. Unlike approaches that retain insights only as inference-time context or retrieved memory, OEL consolidates extracted knowledge into model weights.
- The discussion includes Reflexion and ExpeL as examples of learning from interaction histories.
- It also covers early interaction experience and strategy discovery through reasoning, self-play, and reflection.
6 Conclusion
The conclusion presents OEL as a reward-free extraction-and-consolidation loop that requires no server-side access to user environments. The experiments support iterative gains, shorter responses, and better IF-Eval retention than off-policy consolidation in the tested settings.
- Using real-world deployment as a source of continuous learning remains the broader motivation; the demonstrated evidence comes from text-based games.
Appendix
- A Implementation of On-Policy Context Distillation expands the loss into vocabulary-level reverse KL terms evaluated on student-generated token histories. The full vocabulary sum is approximated using the top-k tokens ranked by student probability, with k = 256 throughout the experiments.
- B Self-Trajectory Variant extracts knowledge independently from each trajectory and pairs it only with prefixes from that trajectory. Algorithm 2 describes this alternative; the experiments instead use the Cross-Trajectory variant with separate extraction and consolidation trajectory sets.
- C Details of Experiments and C.1 Dataset Details specify the 3 × 3 Frozen Lake grid with two holes and the 6 × 6 Sokoban grid with one box. Initial prompts remove explicit mechanics but retain a general objective and movement interface, while subsequent observations describe action outcomes and updated maps.
- C.2 Prompt Templates provides structured and unstructured extraction wrappers using latest_experience and previous_experience, plus a wrapper for applying accumulated experience to new problems. The extraction prompts request broadly applicable rules and strategies rather than board-specific descriptions; the application wrapper warns that inferred experience may be incomplete or incorrect.
- C.3 Extraction Stage and C.4 Consolidation Stage document fixed accumulation steps, K = 10 knowledge samples, length truncation, model-specific configurations, sampling temperature 0.7, and final-checkpoint evaluation. They also disclose that Qwen3-1.7B receives no prior accumulated knowledge during extraction, unlike the other configurations.
- D Experiential Knowledge Examples shows Sokoban knowledge generated by Qwen3-4B-Instruct-2507, including advice about current-state validation, goal tracking, and Manhattan-distance reduction. Some entries claim that direct movement toward the goal is universally valid and describe reaching the goal rather than pushing a box onto it; these are model-generated hypotheses, not verified game rules.
Brief Thoughts
OEL’s main contribution is separating environment-dependent collection from environment-independent on-policy consolidation. The raw-trajectory ablation supports the value of an intermediate knowledge representation, while the self-derived-knowledge comparison shows that a larger source model is not automatically a better teacher for this procedure.
The main limitation is the gap between the deployment motivation and evaluation scope: two small games, short interactions, a few learning rounds, and IF-Eval as the only out-of-distribution test. Different training budgets complicate model-size comparisons, and the erroneous rules in the knowledge examples show that improved pass rates do not establish the correctness of extracted knowledge. The paper does not evaluate long-term distribution shifts, broader capability retention, or the total computational cost of repeated extraction and consolidation. The results demonstrate a learning architecture under controlled conditions, not unrestricted real-world continual learning.