Monthly AI Papers: February 2026
28 Feb 2026 | Paper Review Monthly Papers LLM Agents Agent Training Data Reinforcement LearningContents
Trends
Agent training is increasingly organized around interaction with environments. GLM-5 couples asynchronous reinforcement learning with executable software and terminal tasks; Terminal-Task-Gen builds multi-turn demonstrations from adapted prompts and synthesized tasks; Experiential Reinforcement Learning uses feedback-conditioned reflection and retries. These studies make environment construction and trajectory handling central parts of the training recipe. 1, 2, 3
Learning from unsuccessful attempts takes different forms across the papers. GLM-5 retains erroneous segments as context while masking their SFT loss. Terminal-Corpus performs better when its full synthetic trajectory pool is retained, although filtering also reduces sample count substantially. ERL explicitly generates corrective guidance and distills positive-reward retries into a policy that does not require reflection at deployment. The evidence supports testing recovery-oriented training, but does not establish that failed traces alone cause the gains.
Reported improvements remain sensitive to the success criterion and evaluation conditions. GLM-5’s strong build and individual-check scores coexist with lower complete-project and chained-task performance than Claude Opus 4.5. Nemotron-Terminal improves substantially on Terminal-Bench 2.0 but retains zero scores in several categories. ERL’s largest gains occur in deterministic control episodes limited to eight actions. Together, these results support specific capability gains rather than general claims of reliable autonomous execution.
Top Papers in February 2026
1. GLM-5: from Vibe Coding to Agentic Engineering
- https://arxiv.org/abs/2602.15763
- GLM-5: from Vibe Coding to Agentic Engineering | Summary
- GLM-5: from Vibe Coding to Agentic Engineering presents an open-weight model for planning, implementing, and revising software across extended interactions. Its 744B-parameter MoE activates 40B parameters; base-model training uses a reported 28.5 trillion tokens and extends context to 200K. The pipeline combines DeepSeek Sparse Attention, sequential Reasoning RL, Agentic RL, General RL, and final on-policy cross-stage distillation. For on-policy cross-stage distillation, see this post.
- The agent-training infrastructure separates rollout generation from optimization. Token-in-Token-out preserves sampled token identities, importance sampling masks excessive policy divergence, and stale-trajectory filtering limits policy lag. Executable feedback comes from over 10,000 SWE environments spanning nine languages, alongside terminal and search tasks. These mechanisms address concrete training failures, but their individual contributions and the efficiency gain over synchronous training are not quantified.
- Reported scores include 77.8 on SWE-bench Verified, 73.3 on SWE-bench Multilingual, and 56.2 on the original Terminal-Bench 2.0 under both tested harnesses. On a modified, verified Terminal-Bench set, scores rise to 60.7 with Terminus-2 and 61.1 with Claude Code; these are not matched-set comparisons with other models’ original results. BrowseComp reaches 75.9 with context management, although the report inconsistently describes that condition as Hierarchical Context Management and discard-all.


2. On Data Engineering for Scaling LLM Terminal Capabilities
- https://arxiv.org/abs/2602.21193
- On Data Engineering for Scaling LLM Terminal Capabilities | Summary
- On Data Engineering for Scaling LLM Terminal Capabilities introduces Terminal-Task-Gen and Terminal-Corpus. Existing math, code, and SWE prompts provide broad coverage; seed-based and compositional skill-based generation add terminal-specific tasks. DeepSeek-V3.2 collects trajectories through Terminus 2. Generated tasks use nine shared domain-specific Docker images and synthesized pytest checks, but lack generated oracle solutions; adapted prompt tasks have no verification tests.
- With the Terminus 2 scaffold fixed, SFT produces Nemotron-Terminal-8B, -14B, and -32B. Their Terminal-Bench 2.0 scores are 13.0 ± 2.2, 20.2 ± 2.7, and 27.4 ± 2.4, compared with corresponding Qwen3 baselines of 2.47 ± 0.5, 4.04 ± 1.3, and 3.37 ± 1.6. Selected larger models have lower means, but some uncertainty ranges overlap, and the strongest evaluated systems remain ahead.
- The Qwen3-8B ablations favor mixed single-stage SFT over adapter-then-synthetic training: 13.03 ± 2.16 versus 10.39 ± 1.71. Keeping all 264,207 synthetic trajectories yields 12.4 ± 2.29, compared with 6.74 ± 2.20 for 104,603 complete-only traces and 5.06 ± 2.11 for 83,448 success-only traces. Retaining unsuccessful attempts may help, but these comparisons change data quantity as well as trajectory content and difficulty.


3. Experiential Reinforcement Learning
- https://arxiv.org/abs/2602.13949
- Experiential Reinforcement Learning | Summary
- Experiential Reinforcement Learning trains an experience–reflection–consolidation loop. After an initial attempt receives reward below 1, the same model generates feedback-conditioned guidance and retries. Reflection receives the retry’s reward, while selective self-distillation learns the revised output from the original task context. The distillation filter requires positive retry reward, not improvement over the initial attempt; deployment requires neither the training-time reflection pass nor feedback from a previous attempt.
- ERL exceeds GRPO-based RLVR in all six task–backbone comparisons. For Qwen3-4B-Instruct-2507, evaluation reward rises from 0.86 to 0.94 on Frozen Lake, 0.45 to 0.56 on HotpotQA, and 0.06 to 0.87 on Sokoban. Olmo-3-7B-Instruct improves from 0.39 to 0.66, 0.47 to 0.50, and 0.04 to 0.20, respectively. The largest gain is an absolute reward increase of 0.81, not an 81% relative increase.
- Removing structured reflection lowers final reward in every setting, even though the control retains two attempts and supplies the previous trajectory with a generic improvement instruction. Memory has mixed effects: removing it leaves Qwen’s HotpotQA and Sokoban rewards unchanged and raises Olmo’s Sokoban reward from 0.20 to 0.24. This supports explicit corrective guidance more consistently than cross-episode memory; no ablation removes internalization alone.


Sources
[1] GLM-5: from Vibe Coding to Agentic Engineering
[2] On Data Engineering for Scaling LLM Terminal Capabilities