Gorio Tech Blog search

How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings | Summary

|

Contents

This article explains the key points of How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings.

  • 2026-04-06 (arXiv), Preprint. Under review.
  • Liu, Yujian, Ji, Jiabao, An, Li, Jaakkola, Tommi, Zhang, Yang, Chang, Shiyu.
  • UC Santa Barbara, MIT CSAIL, MIT-IBM Watson AI Lab
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • The paper evaluates whether reusable skill documents help agents when selection, retrieval, and adaptation are required. It assembles 34,198 skills and evaluates three model–harness pairs on 84 SKILLSBENCH tasks and 89 TERMINAL-BENCH 2.0 tasks, with 3 runs per task.
  • Skill benefits generally shrink as task-specific curated skills become harder to access or are removed. On SKILLSBENCH, Claude Opus 4.6 achieves 55.4% with forced loading of curated skills, 38.4% with retrieval excluding curated skills, and 35.4% without skills; Kimi and Qwen fall slightly below their no-skill baselines when curated skills are excluded.
  • Query-specific refinement uses an exploratory task attempt and self-evaluation to compose task-tailored skills, improving pass rates in 7 of 9 tested model–setting combinations. On TERMINAL-BENCH 2.0, retrieval followed by refinement raises Claude’s pass rate from the 57.7% no-skill baseline to 65.5%, but requires extra inference-time exploration; gains are generally smaller in the setting with lower initial skill coverage.

1 Introduction

The paper questions evaluations that directly provide hand-crafted, narrowly task-specific skills, because these conditions bypass discovering useful knowledge in a large repository. Figure 1 illustrates the concern with a flooding task: usgs-data-download, nws-flood-thresholds, and flood-detection collectively specify the data API, threshold source, and detection procedure, making the supplied skills close to a solution guide.

  • The example asks for stations that experienced flooding during April 1–7, 2025, using michigan_stations.txt. Its skills supply details such as nwis.get_iv(), parameter code 00065, truncation to 43 columns, and resample(‘D’).max().
  • The study separately tests selection of available skills, retrieval from a large collection, and adaptation of partially relevant content.
  • Code is available at https://github.com/UCSB-NLP-Chang/Skill-Usage.

The flooding example shows how curated skills can encode nearly the entire task procedure rather than broadly reusable knowledge. The adjacent plot motivates testing discoverability and adaptation: Claude falls from 51.2% with supplied curated skills to 38.4% when retrieval excludes curated skills, while Kimi and Qwen finish slightly below their no-skill baselines. The overall trend is not strictly monotonic for every model.

A task-specific flooding example and SKILLSBENCH pass rates under supplied curated skills, distractors, retrieval with or without curated skills, and no skills.
A task-specific flooding example and SKILLSBENCH pass rates under supplied curated skills, distractors, retrieval with or without curated skills, and no skills.

Reusable knowledge for LLM agents. Prior work represents reusable knowledge as tools, embodied skill libraries, instruction manuals, workflows, procedural memory, and persistent memory. This paper evaluates the utility of retrieved knowledge in a standardized skill format, complementing research focused on creating or evolving that knowledge.

Agentic skills. The standardized artifact consists of a SKILL.md file containing structured metadata and content, optionally accompanied by helper files. The paper positions its evaluation against work on skill infrastructure, discovery, routing, memory, security, and benchmarks that directly provide curated skills.

Agent self-improvement and test-time adaptation. Skill refinement connects to verbal self-reflection, feedback-driven policy improvement, memory-based learning, and inference-time accumulation of reusable knowledge. The evaluated refinement procedures use exploration and self-evaluation rather than access to the benchmark’s ground-truth verifier.

3 Skill Usage in Realistic Settings

The evaluation progressively introduces three challenges: selecting which available skills to load, retrieving candidates from a large repository, and adapting general-purpose skills that only partially match the task. The experiments first establish a skill collection and search system, then vary these challenges on SKILLSBENCH.

3.1 Skill Collection

Skill metadata comes from skillhub.club and skills.sh, while complete skill folders are downloaded from their original GitHub repositories. Filtering retains MIT- and Apache 2.0-licensed skills, removes entries with empty names or descriptions, and deduplicates by file content, producing 34,198 skills across domains including web development, data engineering, development operations, and scientific computing.

  • The collected artifacts include SKILL.md and accompanying helper files, not only metadata descriptions.
  • The source platforms are https://www.skillhub.club/ and https://skills.sh/.

3.2 Skill Search Engine

Skill index. Each skill has a metadata representation formed from its name and description and a full-content representation from SKILL.md. Qwen3-Embedding-4B supplies dense embeddings, while BM25 supplies sparse keyword matching.

Search methods. Direct search embeds the task description once and retrieves top-k candidates from the metadata index. Agentic search allows iterative query formulation, candidate inspection, and selection, with keyword-only, semantic-only, hybrid metadata-only, and hybrid metadata-plus-content variants.

  • The hybrid variants expose keyword, semantic, and combined search tools to the agent.
  • The content-aware variant incorporates full skill content in addition to metadata, allowing retrieval to use information omitted from names and descriptions.

Results. Retrieval is evaluated using task-averaged Recall@k against SKILLSBENCH’s manually curated skills, with Claude Opus 4.6 and Claude Code conducting agentic search. Agentic semantic search improves Recall@3 from 38.1% to 56.8% over direct semantic search, while content-aware agentic hybrid search reaches 65.5% Recall@5 and 68.3% Recall@10 and becomes the default retrieval method.

  • Agentic keyword search reaches only 26.6% Recall@5, compared with 63.1% for agentic semantic search.
  • Adding full content raises hybrid Recall@5 from 63.5% to 65.5% and Recall@10 from 66.7% to 68.3%, but Recall@3 changes from 57.7% to 57.3%.
  • These recall measurements test recovery of curated artifacts; they do not directly measure whether other retrieved skills could also help solve a task.

Iterative agentic search supplies the largest retrieval improvement: using semantic search agentically raises Recall@3 from 38.1% to 56.8%. Content-aware hybrid search gives the best Recall@5 and Recall@10 at 65.5% and 68.3%, although metadata-only hybrid search is slightly better at Recall@3. The metric measures recovery of known curated skills rather than downstream task success.

Recall@3, Recall@5, and Recall@10 for direct semantic retrieval and four agentic retrieval variants using Claude Opus 4.6 with Claude Code.
Recall@3, Recall@5, and Recall@10 for direct semantic retrieval and four agentic retrieval variants using Claude Opus 4.6 with Claude Code.

3.3 Progressive Evaluation Settings

Settings. Six conditions vary whether loading is forced or autonomous, whether skills are supplied or retrieved, and whether task-specific curated skills are available.

  • Curated + forced load explicitly requires loading all original curated skills; Curated leaves loading decisions to the agent.
  • Curated + distractors retains the curated skills and adds retrieved distractors, keeping 5 skills in total.
  • Retrieved (w/ curated) selects top-5 skills from the collection augmented with curated skills; Retrieved (w/o curated) selects top-5 without those artifacts. No skills provides the baseline.

Models and evaluation. Claude Opus 4.6 uses Claude Code, Kimi K2.5 uses Terminus-2, and Qwen3.5-397B-A17B uses Qwen-Code. Each pair independently performs retrieval, task completion, and refinement; evaluation uses 84 SKILLSBENCH tasks, excludes tasks with known issues, and repeats each condition 3 times per task.

  • The results measure end-to-end model–harness pairs rather than model capabilities under a shared harness.
  • Pass rates use the benchmark’s automated verifiers; skill usage tracks whether a trajectory loads any skill and whether it loads all curated skills.

Results. Skill Selection: Agents fail to select the right skills, even when they are directly available. Claude’s pass rate falls from 55.4% with forced loading to 51.2% with autonomous curated-skill loading and 43.5% after adding distractors; the fraction loading all curated skills falls from 49% in Curated to 31% with distractors.

  • Qwen achieves 41.2%, 31.6%, and 33.7% across those three conditions, showing an overall gap from forced loading but not a strictly monotonic decline.
  • Kimi loads any skill in 86.1% of curated-setting trajectories, compared with Claude’s 62.2%, yet its 38.9% curated pass rate is close to the 38.5% result under forced loading.
  • These loading measurements support a selection bottleneck, but loading more skills does not by itself establish effective use of their content.

Skill Retrieval: Requiring agents to retrieve skills further degrades performance. With curated skills still present in the retrieval pool, pass rates fall to 40.1% for Claude, 33.5% for Kimi, and 26.7% for Qwen. Retrieval adds another source of failure: the best tested search configuration recovers only 65.5% of curated skills at Recall@5.

  • Claude’s any-skill loading rate falls from 62.2% with supplied curated skills to 44.4% with retrieved skills.
  • The decline occurs despite retaining task-specific knowledge in the pool, distinguishing discoverability from repository availability.

Skill Adaptation: Without curated skills, agents struggle to adapt general-purpose skills and approach the no-skill baseline. Claude reaches 38.4%, only 3.0 percentage points above its 35.4% baseline, while its any-skill loading rate falls to 16.3%. Kimi reaches 19.8% versus 21.8% without skills, and Qwen reaches 19.7% versus 20.5%.

  • The authors interpret the below-baseline results as evidence that irrelevant skill guidance can divert effort or mislead an agent.
  • The aggregate results show limited utility in this setting, but do not isolate the cause of each failed trajectory or establish a general model-strength effect.

Summary. The evaluation motivates refining both skill metadata and skill content. Metadata refinement may make useful skills easier to recognize, while content refinement may remove noise and make partially relevant knowledge easier to adapt.

The paired panels report task performance and loading behavior across six conditions. Claude achieves 55.4% with forced loading versus 38.4% with retrieval excluding curated skills, where only 16% of trajectories load any skill. Kimi’s comparatively high loading rates do not ensure better performance, and Qwen’s distractor result exceeds its curated result; the overall degradation is therefore not strictly monotonic for every model.

SKILLSBENCH pass rates and trajectory-level skill loading across force-loaded, curated, distractor, retrieved, and no-skill conditions for three model–harness pairs.
SKILLSBENCH pass rates and trajectory-level skill loading across force-loaded, curated, distractor, retrieved, and no-skill conditions for three model–harness pairs.

4 Narrowing the Gap with Skill Refinement

The refinement study asks whether transforming retrieved skills before final task execution can recover lost performance. It compares query-agnostic and query-specific strategies on both SKILLSBENCH retrieval conditions and on TERMINAL-BENCH 2.0, which has no curated skill set.

4.1 Refinement Strategies

Query-agnostic refinement. Each retrieved skill is improved independently without access to the target task or other retrieved skills, approximating an offline-refined collection without processing all 34,198 artifacts. Anthropic’s skill-creator guides generation of synthetic queries, comparisons of agent outputs with and without the skill, self-evaluation, and revision.

  • This procedure is suitable for offline preprocessing, but cannot tailor content to the target task or compose knowledge across retrieved skills.
  • The experiments refine only the retrieved skills; they do not rebuild and search a fully refined collection.
  • The guiding meta-skill is available at https://github.com/anthropics/skills/tree/main/skills/skill-creator.

Query-specific refinement. The agent reads the target instruction and all retrieved skills, attempts an initial solution, self-evaluates without the ground-truth verifier, and creates a task-tailored set of refined skills. It can merge complementary information, discard irrelevant guidance, and incorporate knowledge acquired during exploration.

  • The skill-creator meta-skill guides metadata and content writing.
  • Refinement requires a full exploratory pass at inference time, adding computation beyond simply loading retrieved skills.
  • The appendix limits refinement to one iteration; the resulting set may contain more or fewer skills than the original retrieved set.

4.2 Results

Query-specific refinement is broadly effective. Table 2 reports improvement in 7 of 9 model–setting combinations: on SKILLSBENCH with curated skills in the pool, Claude rises from 40.1% to 48.2% and Qwen from 26.7% to 30.8%. On all 89 TERMINAL-BENCH 2.0 tasks, pass rates rise from 61.4% to 65.5% for Claude, 50.6% to 56.2% for Kimi, and 44.2% to 49.1% for Qwen.

  • The corresponding no-skill TERMINAL-BENCH 2.0 baselines are 57.7%, 46.6%, and 44.7%.
  • The exceptions are Kimi on SKILLSBENCH with curated skills in the pool, which falls from 33.5% to 26.7%, and Claude without curated skills, which changes from 38.4% to 37.9%.
  • Loading rates increase even in these exceptions: Kimi’s with-curated loading rises from 69.7% to 95.2%, and Claude’s without-curated loading rises from 16.3% to 61.1%. Increased adoption is therefore not sufficient for improved correctness.

Pass and Load columns distinguish correctness from adoption. Query-specific refinement improves all three TERMINAL-BENCH 2.0 pass rates, but Kimi’s SKILLSBENCH with-curated pass rate falls from 33.5% to 26.7% despite loading increasing from 69.7% to 95.2%. Query-agnostic results are mixed, and Kimi’s entries are unavailable because its harness lacks required subagent support.

Pass rates and any-skill loading rates before and after query-specific or query-agnostic refinement on SKILLSBENCH and TERMINAL-BENCH 2.0.
Pass rates and any-skill loading rates before and after query-specific or query-agnostic refinement on SKILLSBENCH and TERMINAL-BENCH 2.0.

Figure 3 illustrates cross-skill composition on a TERMINAL-BENCH 2.0 PyTorch tensor-parallelism task. The unrefined agent loads torch-tensor-parallel but ignores pytorch-research, producing a solution that fails at world_size = 2; refinement combines weight-sharding guidance with custom autograd.Function patterns into tensor-parallel-linear, supporting differentiable all_gather and all_reduce and passing world_size = 1, 2, 4.

  • The task requires ColumnParallelLinear to split by columns and combine through all_gather, and RowParallelLinear to split by rows and combine through all_reduce.
  • The example demonstrates successful synthesis of complementary material, but does not estimate how frequently this mechanism explains refinement gains.

Before refinement, the loaded torch-tensor-parallel skill supplies weight-sharding knowledge but lacks differentiable collective wrappers, and the trajectory fails at world_size = 2. Refinement integrates custom autograd.Function patterns from pytorch-research into tensor-parallel-linear, enabling differentiable all_gather and all_reduce and passing world_size = 1, 2, 4. This illustrates cross-skill composition in one successful case.

Before-and-after skill composition and execution outcomes for a PyTorch tensor-parallel linear-layer task.
Before-and-after skill composition and execution outcomes for a PyTorch tensor-parallel linear-layer task.

Query-agnostic refinement yields smaller gains. Claude improves from 40.1% to 42.0% on SKILLSBENCH with curated skills in the pool and from 61.4% to 63.3% on TERMINAL-BENCH 2.0, but changes from 38.4% to 37.4% without curated skills. Qwen’s results are also mixed: 26.7% to 26.2% with curated skills, 19.7% to 24.6% without them, and 44.2% to 44.9% on TERMINAL-BENCH 2.0.

  • Kimi’s query-agnostic results are omitted because Terminus-2 lacks the subagent operations required by the procedure.
  • Query-specific refinement is not uniformly better: Qwen’s without-curated result is higher under query-agnostic refinement than under query-specific refinement, which reaches 21.5%.

Refinement effectiveness depends on initial skill quality. GPT-5.4 scores the collective relevance and coverage of initially retrieved skills on a 1–5 scale. SKILLSBENCH with curated skills scores 4.01, 3.83, and 3.85 for Claude, Kimi, and Qwen; without curated skills it scores 3.49, 3.31, and 3.39; TERMINAL-BENCH 2.0 scores 4.02, 3.96, and 4.08.

  • The higher-coverage settings generally show larger query-specific gains, consistent with refinement extracting and composing useful information from retrieved material.
  • This is a setting-level association based on an LLM judge, not proof that refinement cannot generate new knowledge or that coverage alone determines success. The refinement prompt explicitly permits adding knowledge discovered during exploration.
  • Kimi’s decline in the high-coverage with-curated setting is an exception to a simple coverage-based explanation.

Initial coverage is lower for SKILLSBENCH without curated skills, with scores from 3.31 to 3.49, than for the with-curated and TERMINAL-BENCH settings, which range from 3.83 to 4.08. This pattern is consistent with refinement benefiting from relevant source material, but the scores are LLM judgments and do not explain all exceptions, including Kimi’s decline in the with-curated condition.

Average GPT-5.4 coverage scores for initially retrieved skill sets across three evaluation settings and three models.
Average GPT-5.4 coverage scores for initially retrieved skill sets across three evaluation settings and three models.

5 Conclusion

The paper concludes that skill utility depends on agents discovering, loading, and using relevant content, not merely on repository availability. Query-specific refinement recovers performance in several settings with relevant retrieved knowledge, while gains are limited in the lower-coverage SKILLSBENCH setting without curated skills. The authors identify retrieval, offline refinement, and compatibility with differing model capabilities as remaining research needs.

Ethics Statement

The Ethics Statement reports that the study uses public benchmarks and publicly available GitHub skills filtered to MIT and Apache 2.0 licenses, without private or sensitive data. Evaluations run in isolated Docker containers, which the authors state avoids risks to external systems.

LLM Usage Disclosure

The LLM Usage Disclosure cites COLM’s policy and reports LLM assistance with modifying evaluation infrastructure, debugging, and analyzing trajectories. Writing assistance included revising author-drafted text, proofreading, plotting scripts, and LaTeX formatting; the authors attribute research ideas, experimental design, and analysis to themselves.

Appendix

  • A Skill Search Engine Details. SQLite FTS5 implements BM25 retrieval with field weights of 10 for name, 5 for description, and 5 for full content when included; dense queries use Qwen3-Embedding-4B with the prefix “Find skills matching this query:”. The /keyword, /semantic, and /hybrid endpoints provide retrieval, while /detail retrieves full SKILL.md content. Hybrid search uses the RRF score (\sum_s \frac{w_s}{k+r_s}), with default method weights 0.5 and k = 60; semantic metadata–content combination uses (1 − w) · simmeta + w · simcontent, with w = 0.05 selected through synthetic-query Recall@5 sweeps. The finding-skills protocol decomposes tasks, searches iteratively, inspects candidates, and records 10 selected skills for retrieval evaluation, while downstream task conditions use top-5 skills.
  • B Experiment Details. SKILLSBENCH excludes mhc-layer-impl, scheduling-email-assistant, and fix-visual-stability, leaving 84 tasks; TERMINAL-BENCH 2.0 uses all 89 tasks. Claude uses Claude Code v2.1.19; Kimi uses Terminus-2 with a maximum input of 253,952 tokens and local SGLang serving; Qwen/Qwen3.5-397B-A17B-FP8 uses Qwen-Code v0.12.3 and local SGLang serving. Harbor runs each task 3 times in isolated Docker containers and evaluates outputs with automated verifiers. SKILLSBENCH uses a 1.5× timeout multiplier for all models; TERMINAL-BENCH 2.0 uses 2× for Kimi and Qwen and 1× for Claude.
  • C Skill Refinement Details. C.1 Query-Specific Refinement operates in the task’s own Docker environment with its data, libraries, and tools, but without the ground-truth verifier, and permits one iteration of exploration and revision. Its procedure reads all retrieved skills, tests their guidance during a task attempt, reflects on useful and misleading content, and writes composed task-appropriate skills, including knowledge discovered during exploration. C.2 Query-Agnostic Refinement processes each skill independently in Ubuntu 24.04 with Python, uses skill-creator-guided synthetic queries and A/B testing, and likewise permits only one improvement iteration.
  • D LLM-as-Judge for Skill Coverage. GPT-5.4 receives the task instruction and complete retrieved skill material, including helper files, truncated to 400K characters per skill and 2M characters overall. It rates collective coverage from 1, meaning no relevant coverage, to 5, meaning complete coverage, with intermediate levels distinguishing low, moderate, and high coverage. Its output records a score, covered material, and remaining gaps; these judgments supply Table 3 rather than the benchmark pass/fail results.

Brief Thoughts

The progressive evaluation design is the paper’s main contribution: retaining curated skills while adding distractors or requiring retrieval exposes different failure points before curated knowledge is removed entirely. Reporting loading behavior alongside verifier-based pass rates also reveals cases where increased skill adoption accompanies worse task performance, as with Kimi after query-specific refinement in the with-curated setting.

Query-specific refinement adds a full exploratory attempt, but the paper does not report a compute-matched extra-attempt baseline, detailed runtime or token costs, or uncertainty intervals for the pass rates. Different harnesses and timeout settings limit model-only comparisons, while LLM-judged coverage supports an association rather than a causal explanation. The experiments establish limitations of retrieved skills in the tested coding and terminal tasks, not a universal failure of reusable skills.