SkillOS: Learning Skill Curation for Self-Evolving Agents | Summary
07 May 2026 | Paper Review Agentic Skills Self-Evolving Agents Reinforcement LearningContents
- Summary
- 1. Introduction
- 2. Related Work
- 3. Methodology
- 4. Experiments
- 5. Analysis
- 6. Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of SkillOS: Learning Skill Curation for Self-Evolving Agents.
- 2026-05-07 (arXiv)
- Ouyang, Siru, Yan, Jun, Chen, Yanfei, Han, Rujun, Wang, Zifeng, Mishra, Bhavana Dalvi, Meng, Rui, Li, Chun-Liang, Jiao, Yizhu, Zha, Kaiwen, et al.
- University of Illinois Urbana-Champaign, Google Cloud AI Research, Massachusetts Institute of Technology
- Paper
Summary
- SkillOS: Learning Skill Curation for Self-Evolving Agents studies how agents convert past task executions into reusable procedural knowledge. It separates a frozen Agent Executor from a trainable Skill Curator that inserts, updates, and deletes Markdown skills in an external SkillRepo.
- The curator is trained with GRPO on streams of related tasks, allowing skills extracted from earlier trajectories to be evaluated through later task outcomes. Its composite reward combines judge-derived downstream success, valid function calls, judged content quality, and repository compression.
- With Qwen3-8B as executor, SkillOS raises reported ALFWorld success from 47.9% without memory to 61.2%, while reducing average actions from 21.1 to 18.9. The trained 8B curator also improves larger, unseen executors, but cross-domain transfer has exceptions, and action savings do not establish lower total system cost.
1. Introduction
Streaming agents must solve each task before observing future tasks, yet one-off execution discards experience that could improve subsequent decisions. The paper identifies skill curation—not merely storage or retrieval—as the central problem: extracting useful lessons, revising existing procedures, and removing low-utility knowledge.
- Manual repositories require expert effort, while fixed prompting and heuristic management do not directly optimize for the executor’s downstream needs.
- Earlier RL approaches often train skill use or short-stream adaptation. SkillOS instead evaluates curation through related future tasks and supplements delayed outcome feedback with intermediate rewards.
2. Related Work
Memory for Self-Evolving Agents distinguishes case-based memory, such as raw trajectories and query–response exemplars, from strategy-based memory, such as workflows, distilled insights, and reusable skills. SkillOS follows the modular skill paradigm but simplifies each skill to one Markdown file rather than a folder containing scripts and supporting resources.
- Learning Memory and Skill Curation with RL distinguishes context management, memory retrieval, skill utilization, and skill curation as optimization targets.
- SkillRL and D2Skill train models to use externally curated skills, while ARISE trains retrieval and execution with heuristic skill management.
- SkillOS targets curation itself, using related task streams to provide opportunities for learning revision and deletion beyond immediate post-insertion transfer.
3. Methodology
The methodology combines a modular execution–curation loop for streaming inference with an RL procedure for training the curator. Executor parameters remain fixed; inference changes the external repository, while training changes the curator’s policy for managing it.
3.1. Streaming Skill Curation with Multi-Agent Modular Design
For each arriving task, BM25 retrieves relevant skills from SkillRepo, and the frozen executor conditions its actions on the task, current observation, and retrieved content. After execution, the curator receives the trajectory, a self-judged correctness signal, and related skills, then produces structured insert_skill, update_skill, or delete_skill operations.
- Each skill contains YAML frontmatter with name and description, followed by Markdown instructions encoding workflows, constraints, and reusable heuristics.
- Applying the operations updates SkillRepo for subsequent tasks, allowing adaptation through procedural memory without changing executor weights.
- The appendix function signatures use new_skill_insert, skill_update, and skill_delete, whereas the main formulation names the operations insert_skill, update_skill, and delete_skill.
The diagram locates adaptation in the external repository: executor weights stay frozen, while experience reaches the curator and changes stored skills. Its Markdown example distinguishes suggested sections from additional structure that can emerge during training.
3.2. Learning Skill Curation with RL
3.2.1. Training Instance Construction builds each training instance as a sequential group of skill-related tasks. Each rollout starts with an empty SkillRepo; trajectories from earlier tasks update it, and later tasks test whether those updates are useful.
- For reasoning problems, Gemini-2.5-Pro annotates topics, capabilities, concepts, strategies, and pitfalls; semantic similarity and dependency filters then construct groups.
- For ALFWorld and WebShop, the experiments use default task-type annotations rather than the full reasoning annotation pipeline.
- Grouping creates repeated opportunities for skill reuse and refinement rather than testing only immediate transfer after one memory update.
3.2.2. Training Loop and Policy Optimization uses the composite reward r = r^task + λ_f r^fc + λ_u r^cnt + λ_c r^comp. The coefficients are λ_f = 1.0, λ_u = 0.1, and λ_c = 0.05.
- Task outcome reward averages success over tasks 2 through |G|, excluding the first because it runs before any curator update. Correctness signals are obtained with an LLM judge using the frozen executor backbone, rather than the final benchmark scoring procedure.
- Function call reward averages the fraction of valid, successfully executed operations; content quality reward uses Qwen3-32B to judge whether generated skills are meaningful and reusable.
-
Compression reward averages (1-\frac{ \mathcal{S}_i }{ \chi_i }), where the numerator is repository token length after an update and the denominator is curator input-context token length. It discourages copying complete trajectories into memory.
For each task group, training samples multiple complete curation rollouts, each with its own evolving repository and executor trajectories. GRPO uses each rollout’s reward minus the mean reward of the sampled rollouts as its advantage, then applies a clipped importance-ratio objective.
- The same sequence-level advantage is assigned uniformly to every generated curation token; the auxiliary rewards therefore add supervision without providing operation-specific causal credit.
- The main text states that the KL term is discarded to encourage exploration, although Table 4 lists a KL loss coefficient of 0.001.
Each rollout begins with an empty repository and processes a related task group, allowing earlier edits to affect later execution. Four reward signals supervise the curator; the lower schematic depicts the intended progression toward more varied management operations, rather than measured evidence that every operation is mastered.
4. Experiments
Experiments cover multi-turn interaction and single-turn reasoning, followed by transfer tests across executor backbones and task domains. The evaluation separates task effectiveness from executor action efficiency.
4.1. Setup
Agentic evaluation uses ALFWorld household tasks and WebShop shopping instructions; reasoning evaluation uses AIME24, AIME25, and GPQA-Diamond. Reasoning training samples 33,000 problems from DeepMath-103k, while agentic training uses the corresponding benchmark training splits.
- Both the curator and frozen training executor are Qwen3-8B. GRPO uses learning rate 1 × 10−6, batch size 32, and 8 sampled rollouts per group.
- Training uses verl on 16 H100 GPUs and takes approximately 3 days for ALFWorld, 2.5 days for reasoning, and 5 days for WebShop.
- Inference retrieves the top 5 skills, permits 30 turns for agentic tasks, and retains 3 recent interaction steps. ReAct is used for agentic execution and CoT for reasoning.
- Comparisons include No Memory, ReasoningBank, MemP, untrained SkillOS-base, and Gemini-2.5-Pro-based SkillOS-gemini. Reported means and standard deviations come from 3 runs.
4.2. Main Results
On ALFWorld, the trained curator achieves the highest reported average success for all three main executors: 61.2 ± 4.6% with Qwen3-8B, 68.6 ± 5.7% with Qwen3-32B, and 80.2 ± 3.1% with Gemini-2.5-Pro. The corresponding SkillOS-base results are 53.1%, 59.8%, and 70.7%, supporting a benefit from curator training under the same repository interface.
- For Qwen3-8B, ReasoningBank reaches 55.7% success with 20.1 actions, compared with SkillOS at 61.2% and 18.9 actions.
- SkillOS-gemini reaches 50.7%, 63.6%, and 79.3% across the three executors. The trained 8B curator exceeds direct Gemini-2.5-Pro curation on the aggregate metric, although it does not win every task subset across all executor blocks.
- Several configurations show substantial run-to-run variation; the paper does not report significance tests.
Across all three executor blocks, SkillOS has the highest reported average ALFWorld success and the fewest actions. Per-category results show that this aggregate advantage does not imply winning every subset, and several configurations have substantial run-to-run variability.
On WebShop, SkillOS obtains scores of 40.6, 49.2, and 56.0 with Qwen3-8B, Qwen3-32B, and Gemini-2.5-Pro, respectively; success rates are 16.5%, 16.5%, and 41.3%. Average reasoning accuracy across AIME24, AIME25, and GPQA-Diamond rises from the No Memory results of 69.6%, 74.0%, and 81.8% to 73.8%, 79.7%, and 88.6%.
- With Gemini-2.5-Pro as executor, SkillOS achieves 92.2% on AIME24, 86.7% on AIME25, and 86.8% on GPQA-Diamond.
- Some untrained or heuristic memory configurations underperform No Memory, so retaining experience alone does not ensure improvement.
- The baseline curator backbone varies by table: ALFWorld uses Qwen3-8B for ReasoningBank and MemP across executor blocks, whereas WebShop and reasoning use the executor’s backbone for those methods.
ALFWorld action counts fall from 21.1 to 18.9, 20.3 to 17.3, and 17.7 to 14.8 relative to No Memory across the three executors. WebShop also uses fewer actions than No Memory, but SkillOS-gemini is more action-efficient with the Gemini-2.5-Pro executor: 17.8 actions versus 18.3 for SkillOS.
- The authors hypothesize that larger agentic gains reflect reusable action ordering, exploration, recovery procedures, and environment constraints; reasoning skills instead encode more abstract decomposition and verification strategies.
- The qualitative examples illustrate this proposed explanation but do not isolate the mechanism experimentally.
- Although the setup identifies reasoning-token count as an efficiency metric, Table 2 reports reasoning accuracy rather than token counts. The action comparisons exclude curator inference, judging, and training costs.
The trained curator improves WebShop effectiveness and average reasoning accuracy with all three executors. Action efficiency is not uniformly best: in the Gemini-2.5-Pro WebShop block, SkillOS-gemini uses 17.8 actions compared with 18.3 for SkillOS; reasoning-token efficiency is not reported here.
4.3. Generalization of SkillOS
The curator trained with a Qwen3-8B executor transfers to Qwen3-32B and Gemini-2.5-Pro without retraining those executors. On ALFWorld, it improves Gemini-2.5-Pro from 66.4% without memory to 80.2%, showing useful transfer to a stronger backbone.
- The authors interpret weaker results from direct Gemini-2.5-Pro curation, especially with the small executor, as possible curator–executor mismatch.
- The experiments do not independently measure that mismatch or separate it from other differences in generated skill content.
Cross-domain tests transfer the learned curator, not simply a fixed collection of source-domain skills. Most combinations improve over No Memory, but the heatmaps contain negative entries, so transfer is not uniformly beneficial.
- With Gemini-2.5-Pro, the reasoning-trained curator improves ALFWorld by 7.2 points and WebShop score by 0.7; with Qwen3-32B, reasoning-to-WebShop transfer instead reduces score by 1.2.
- WebShop-to-reasoning transfer changes accuracy by −1.2, −2.8, and +0.3 points for Qwen3-8B, Qwen3-32B, and Gemini-2.5-Pro.
- The caption calls the changes relative improvements, but the displayed values are absolute differences from the baseline metrics, not percentage-normalized gains.
Rows identify the curator's training domain and columns its evaluation domain. Most off-diagonal changes are positive, but negative entries—such as WebShop-to-reasoning with Qwen3-32B at −2.8—show that learned curation does not transfer beneficially in every pairing.
5. Analysis
ALFWorld ablations use Qwen3-8B for both roles. Removing content-quality reward lowers success from 61.2% to 58.6%, removing compression reward lowers it to 60.0%, and replacing grouped streams with random sequences lowers it to 57.3%.
- Average actions also increase from 18.9 to 20.1, 19.3, and 20.6, respectively.
- Removing grouping produces the largest observed degradation among these ablations, supporting related future tasks as supervision for curation.
- Table 3 does not ablate task outcome or function-call reward and provides no standard deviations for the ablated configurations.
Removing task grouping causes the largest success-rate drop among the tested ablations, from 61.2 to 57.3. Removing content-quality or compression reward also reduces success and increases actions, supporting these design choices without establishing the contribution of every reward component.
The operation distribution shifts from insertion-dominated behavior toward more frequent updates as training proceeds. Deletion remains a small fraction throughout, making consolidation and refinement the clearest behavioral change.
- Operation frequencies show changing management behavior but do not establish whether individual updates or deletions improve future outcomes.
- The low deletion share leaves limited evidence for learning deletion as a major long-term curation capability.
Insertion declines while updates become more frequent over training, indicating a shift toward refining existing entries. Deletion remains uncommon, so the plot provides stronger evidence for consolidation than for extensive learned forgetting.
Repository analysis tracks new Markdown sections and changing skill categories. Early section additions emphasize generic guidance, whereas later additions more often include failure handling and conditional branches; task-specific entries shrink while the repository contains broader strategies for verification, search, recovery, and adjustment.
- These observations suggest that training changes the structure of stored procedures as well as their quantity.
- Figure 5 also identifies state-verification skills as dominant early on, so its category trends should not be reduced to a repository initially composed only of narrow task-specific skills.
- The analysis is descriptive; it does not isolate the downstream contribution of particular emergent sections or meta-skills.
The left panel tracks the proportions of skills containing newly introduced section types, with later additions including conditional branches and failure handling. The right panel tracks category counts, showing shrinking task-specific content and diversified strategies; its early state-verification dominance complicates the text's simpler task-specific-to-meta-skill narrative.
On ALFWorld, SkillOS raises explicit skill usage from 87.9% to 100.0%, success among skill-using examples from 53.6% to 61.2%, and repository coverage from 72.9% to 88.6% relative to SkillOS-base. Average skills used per example fall from 2.24 to 1.95.
- Broader repository coverage together with fewer skills per example is consistent with more targeted use.
- These statistics associate skill use with success; they do not causally attribute outcomes by removing or replacing individual retrieved skills.
6. Conclusion
The conclusion identifies modular, executor-grounded curation as the principal contribution: a small trained curator improves a frozen executor through an evolving procedural repository. Grouped streams and composite rewards provide supervision for this policy, with benefits demonstrated on the evaluated benchmarks and several executor backbones.
Appendix
- A. Prompts documents the system’s templates. A.1. Prompt for Skill Curator supplies the task, previous skills, trajectory, and result, and requests atomic, reusable Markdown skills with name and description frontmatter; A.2. Prompt for Agent Executor supplies retrieved skills and recent interaction context for agentic tasks, or the problem and retrieved skills for reasoning, with explicit step-by-step reasoning instructions.
- A.3. Prompt Used During Training asks the Qwen3-32B content judge to assess abstraction, reusability, actionability, and faithfulness. A.4. Prompt for LLM-as-a-Judge to Obtain Correctness Signals uses the corresponding frozen executor to evaluate trajectories or solutions; the WebShop judge defines success as score >= 0.5, whereas final benchmark success requires the environment reward to equal exactly 1.
- B. Implementation Details and B.1. Hyperparameters specify maximum prompt and response lengths of 16,384 and 4,096, temperature 1.0, and 60, 50, and 100 training steps for ALFWorld, WebShop, and reasoning. Task-stream grouping sizes are 10 for each agentic benchmark and Random(5,12) for reasoning, distinct from the GRPO rollout group size of 8; Table 4 also lists the KL coefficient of 0.001 that conflicts with the main-text description.
- B.2. Grouping Training Instances uses B.2.1. Stage 1: Latent Attribute Annotation to describe reasoning problems along five dimensions: topics, skills, concepts, heuristic strategies, and pitfalls. Gemini-2.5-Pro produces standardized short phrases through structured decoding; B.2.2. Stage 2: Group Construction matches them with all-MiniLM-L6-v2 and soft-Jaccard similarity, then requires shared concepts and skills, shared strategies or pitfalls, nonredundancy, a new concept or skill, and nondecreasing difficulty.
- The grouping pipeline uses cosine threshold 0.60, at least 1 matched concept pair and 1 matched skill pair, maximum topic soft-Jaccard 0.65, and overall similarity between 0.30 and 0.85. Dimension weights are (5, 4, 3, 1, 2) for concepts, skills, strategies, pitfalls, and topics; curriculum probabilities are (0.80, 0.20, 0.00), with easy→hard gaps [0.5, 3.0] and same-mode maximum gap 0.3. Candidates come from a dependency-phrase inverted index capped at 2,000 or a fallback pool of 200, and admissible successors are ranked by weighted similarity plus a difficulty bonus with weight 1.0; the weights were tuned by inspecting pairs for 200 held-out source tasks.
- B.3. Experiment Setup expands B.3.1. Datasets, B.3.2. Baselines, and B.3.3. Evaluation Metrics. ALFWorld uses 3,553 training tasks and 140 valid_seen evaluation tasks; WebShop uses 10,587 training instructions and 500 test instructions; AIME24 and AIME25 contain 30 problems each, and GPQA-Diamond contains 198. The baselines span independent memory-free execution, ReasoningBank’s distilled insights, MemP’s heuristic procedural-memory management, and the same SkillOS interface with either an untrained Qwen3-8B curator or direct Gemini-2.5-Pro curation.
- Reasoning preprocessing starts with approximately 33,000 annotated problems and reports a final set of 20,000 grouped training instances. Evaluation uses environment success for agentic tasks, partial-match WebShop score, actions averaged over successful and failed episodes, mathematical-equivalence verification for AIME through https://github.com/huggingface/Math-Verify, and exact option matching for GPQA. The appendix defines ALFWorld Avg. SR as a macro-average across six categories, but the displayed table values correspond to task-count-weighted averages: for example, the Qwen3-8B No Memory category rates yield approximately 45.3% unweighted rather than the reported 47.9%.
- C. Additional Analyses includes C.1. Results on Gemini-3.1-Flash-Lite: SkillOS reaches 73.1 ± 2.7% ALFWorld success with 15.5 actions, compared with ReasoningBank at 66.0 ± 2.7% and 17.6 actions, and No Memory at 61.2 ± 2.3% and 18.5 actions. C.2. Case Studies illustrates failure recovery that references other skills, alternative mathematical solution paths with prerequisites, more concrete counting guidance than SkillOS-base, and successful CD inspection under a desklamp after the memory-free agent exhausts its 30-step budget.
- D. Limitations identifies three deliberate simplifications: keyword-based BM25 retrieval, single-file declarative Markdown skills without supporting executable resources or an explicit hierarchy, and a frozen executor. These choices isolate curation but leave retrieval quality, richer skill representations, and curator–executor co-adaptation unresolved.
- E. Future Research Directions proposes agentic search over experiential memory, hierarchical and compositional skills, and shared memory across agents with conflicting updates and difficult credit assignment. F. Use of LLMs states that LLMs assisted with writing clarity and readability, while the authors retained responsibility for ideas, methodology, experiments, analyses, and final writing decisions.
Brief Thoughts
The SkillOS-base comparison under the same architecture and memory interface, together with the grouped-stream ablation, provides the clearest evidence for the training procedure. These comparisons support learning the curation policy rather than attributing the gains solely to external memory or a larger curator.
The results are less conclusive about long-term scalability and precise credit assignment: training streams contain 10 agentic tasks or 5–12 reasoning tasks, deletion remains rare, and all curation tokens share one rollout advantage. Negative transfer entries, unreported end-to-end costs, small reasoning evaluation sets, and inconsistent KL and ALFWorld averaging descriptions limit claims of uniformly efficient, robust lifelong self-evolution.