Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills | Summary
18 Dec 2025 | Paper Review Agent Learning Agent's Memory Agentic SkillsContents
- Summary
- 1 Introduction
- 2 Background
- 3 Overview of Adaptation Paradigms of Agentic AI
- 4 Agent Adaptation
- 5 Tool Adaptation
- 6 Comparison of Adaptation Paradigms
- 7 Evaluation
- 8 Applications
- 9 Opportunities
- 10 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills.
- 2025-12-18 (arXiv)
- Jiang, Pengcheng, Lin, Jiacheng, Shi, Zhiyi, Wang, Zifeng, He, Luxi, Wu, Yichen, Zhong, Ming, Song, Peiyang, Zhang, Qizheng, Wang, Heng, et al.
- UIUC, Stanford, Princeton, Harvard, UC Berkeley, Caltech, UW, UCSD, Georgia Tech, Northwestern, TAMU, MGH, Keiji AI, Unity
- Paper
- Github
Summary
- Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills reviews literature published through early 2026 according to what changes in an agentic system and where its supervision originates. The four paradigms distinguish agent adaptation using tool-execution feedback (A1), agent adaptation using evaluated outputs (A2), independently trained tools (T1), and tools adapted under fixed-agent supervision (T2).
- The survey connects post-training with persistent memory and reusable procedural skills, covering model updates, external memory curation, and specialized subagents. Its retrieval case study reports s3 at 58.9% average accuracy with 2.4k training samples, compared with roughly 170k examples for Search-R1, but differences in backbones and architecture prevent attributing the efficiency gap solely to paradigm choice.
- The evaluation framework combines execution quality, final task utility, learning dynamics, cost, safety, and alignment. Controlled cross-paradigm comparisons, longitudinal memory and skill evaluation, and stable agent-tool co-adaptation remain open problems; the paper contributes a literature synthesis rather than a new trained system or controlled experimental study.
1 Introduction
The introduction identifies unreliable tool use, weak long-horizon planning, domain reasoning gaps, and poor transfer to unfamiliar environments as motivations for adaptation. It connects post-training, memory, and skills within a classification of agent and tool adaptation; Figure 1 summarizes the classification, and Figure 2 maps the paper’s structure.
- Methods are assigned by their dominant optimization target and signal source, with overlapping and hybrid cases identified explicitly.
- The stated scope extends through early 2026 and draws primarily on major machine learning and NLP venues and arXiv preprints. The methodology paragraph does not specify search queries, screening procedures, or an inclusion count.
The diagram separates changes to the agent from changes to callable components and traces their supervision routes. Placing memory and skills on the tool side brings external accumulation of experience into the same framework as model training.
The hierarchy develops the taxonomy before method catalogs, comparisons, evaluation, applications, and opportunities. Memory and skills appear within T2, while evaluation examines the paradigms across separate dimensions.
2 Background
The background defines the components of an agentic system and distinguishes prompt-based adaptation from fine-tuning. These approaches change the model’s input context and internal parameters, respectively.
2.1 Agentic AI Systems
An agentic AI system combines a foundation model with planning, tool use, and memory to execute multi-step tasks using environmental feedback. The stated focus is single-agent adaptation, although later examples include multi-agent coordination.
- Static planning decomposes tasks into reasoning paths; dynamic planning, illustrated by ReAct and Reflexion, revises plans using observations and past outcomes.
- Tool selection and interpretation of tool outputs are adaptation surfaces. Short-term memory supports the current task, while long-term memory retains knowledge and experience across sessions.
- Memory design involves representation, retention, retrieval, and integration into reasoning. The survey treats most external adaptive memory as a tool, while distinguishing parametric and hybrid cases.
2.2 Adaptation
Prompt Engineering changes instructions, examples, goals, or constraints without updating model weights. Fine-Tuning updates parameters through full training or parameter-efficient methods such as LoRA; SFT imitates demonstrations, DPO uses preference signals, and PPO or GRPO optimizes behavior through rewarded interaction.
- Prompt adaptation avoids additional model training; parameter updates can internalize task-specific behavior but require training resources.
- The paper characterizes LoRA as typically matching full fine-tuning quality while updating less than 1% of parameters. This is a cited generalization, not a result established by the survey.
- RL can learn from rewards without complete demonstrations, but the paper describes it as harder to stabilize than SFT.
3 Overview of Adaptation Paradigms of Agentic AI
The framework classifies methods by the component being adapted and the source of supervision. Figure 3 traces the optimized components and feedback routes for SFT and RL.
3.1 Mathematical Notations
The notation separates adaptation targets, data sources, and objectives: the agent and external tools; offline data and interaction environments; and measures of performance or alignment. Offline data can contain demonstrations, synthetic trajectories, or interaction logs, while environments provide online feedback.
- The agent is the foundation model responsible for reasoning and decisions. Tools include retrievers, planners, executors, simulators, and dynamic external memory.
- Objectives include imitation losses for offline learning and outcome measures such as task success for environmental interaction.
- Parametric and hybrid memory can cross the agent-tool boundary; classification therefore depends on which parameters or external stores change.
3.2 Four Adaptation Paradigms of Agentic AI
A1 optimizes the agent using evaluated tool-execution results; A2 optimizes the agent using its final output, with or without tools. T1 trains tools independently of a particular agent, while T2 adapts tools for the downstream utility of a fixed agent.
- A1 SFT imitates successful reference actions, and A1 RL rewards execution outcomes. A2 SFT can combine tool-call imitation with final-answer supervision because answer supervision alone need not teach tool invocation.
- T2 supervision can weight trajectories by final-output quality, derive tool targets from output consistency, or reward the frozen agent’s final task performance.
- External memory updated from frozen-agent outputs is predominantly T2; independently trained memory components are predominantly T1. Changes to the agent’s own parameters belong under A1 or A2, while auxiliary parameter updates create boundary cases.
The highlighted paths distinguish optimization based on tool results from optimization based on final agent outputs. T2 routes supervision through the fixed agent to the tool; T1 training is independent of the consuming host.
3.3 Illustrative Examples
Paired retrieval and code examples show that the same tool interaction can support different paradigms depending on where the reward is computed. DeepRetrieval rewards retrieval quality, Search-R1 evaluates the final answer, and ReTool incorporates executed code as context before rewarding the final response.
- DeepSeek-R1 code training is classified as A1 when the reward derives solely from sandbox test-case pass rates; the classification depends on that condition.
- A contrastively trained dense retriever is T1 when its training is independent of the consuming agent. A trained query-rewriting agent can also serve as T1 after being frozen.
- s3 and AgentFlow illustrate T2: search or planning components learn from a fixed generator’s downstream output quality.
4 Agent Adaptation
Agent adaptation can target execution mechanics or overall reasoning and coordination. A1 feedback identifies execution failures, while A2 rewards task outcomes but makes attribution across intermediate decisions more difficult; the density of either signal depends on its implementation.
4.1 A1: Tool Execution Result as Signal
A1 methods use evaluated tool outcomes as supervision, although successful individual calls do not establish that their composition solves a task. Figure 4 and Table 1 organize SFT, off-policy preference learning, and RLVR methods across retrieval, code, formal proofs, routing, SQL, recommendation, and OCR.
- Earlier Works: SFT & Off-Policy Methods distinguishes golden-answer, golden-format, and direct-execution alignment. Toolformer selects calls through likelihood improvement; ToolAlpaca retains successful trajectories, TRICE ranks execution-assisted responses, and TP-LLaMA uses failed branches as preference negatives.
- Gorilla uses AST matching for call structure, while ToolFlow constructs coherent multi-tool training data. Structural validity can admit semantically wrong calls; CodeAct, NExT, AutoTools, LeReT, and RetPO use execution information for code generation, repair, tool construction, or query rewriting.
- RLVR-Based Methods emphasizes reward density and composition. DeepRetrieval reports recall of 65.1% versus 24.7%; RLEF uses multi-turn public-test feedback, Code-R1 prioritizes reliable reward pipelines, and proof assistants provide verifiable proof-state transitions.
- The reviewed multi-tool methods address routing, environment construction, and step-level feedback. The synthesis emphasizes reward quality, task and format rewards, KL regularization, curricula, and sampling, while recognizing rollout and execution costs.
The chronology organizes early tool-use training and later execution-reward methods across code, retrieval, proofs, routing, and OCR. It provides historical placement without comparable performance or adoption measurements.
The catalog separates SFT and off-policy methods from RLVR and records tasks, tools, backbones, and tuning procedures. Its heterogeneous rows describe A1 implementations rather than a common experimental comparison.
The catalog separates SFT and off-policy methods from RLVR and records tasks, tools, backbones, and tuning procedures. Its heterogeneous rows describe A1 implementations rather than a common experimental comparison.
The catalog separates SFT and off-policy methods from RLVR and records tasks, tools, backbones, and tuning procedures. Its heterogeneous rows describe A1 implementations rather than a common experimental comparison.
4.2 A2: Agent Output as Signal
A2 adapts reasoning and tool-use strategy through evaluations of the agent’s outputs. Figure 5 and Table 2 distinguish methods without tools from methods that additionally learn when to invoke tools and how to incorporate their results.
- Agent Adaptation w/o Tools covers scalar-reward RL, inference-time refinement, and linguistic feedback. Self-Refine revises outputs without weight changes; SCoRe trains self-correction, and EHRMind uses SFT warm-up before RL.
- TextGrad uses natural-language critiques to optimize black-box systems. The survey reports GPT-4o accuracy increasing from 26% to 36% on LEETCODE-HARD and from 91.2% to 95.1% on MMLU-Physics; metaTextGrad refines the optimizer’s prompts.
- Agent Adaptation w/ Tools covers retrieval, code-assisted reasoning, curricula, reflection, and infrastructure. ReSearch integrates <think>/<search>/<result> tags, ReTool incorporates code execution into rollouts, and Agent Lightning, A²FM, and VerlTool address training or execution infrastructure.
- Sparse final-output rewards complicate credit assignment. Search-R1’s roughly 170k examples illustrate one training scale, while reviewed systems report self-correction, query refinement, and adaptive search depth.
The timeline groups inference-time refinement, reasoning RL, search learning, and training infrastructure under output-signaled adaptation. Methods differ in whether weights change despite sharing the survey’s signal-based classification.
Grouping methods with and without tools exposes the additional coordination problem introduced by tool use. The tuning column shows that A2 encompasses prompt optimization and test-time changes alongside SFT and RL.
Grouping methods with and without tools exposes the additional coordination problem introduced by tool use. The tuning column shows that A2 encompasses prompt optimization and test-time changes alongside SFT and RL.
Grouping methods with and without tools exposes the additional coordination problem introduced by tool use. The tuning column shows that A2 encompasses prompt optimization and test-time changes alongside SFT and RL.
5 Tool Adaptation
Tool adaptation changes external components while retaining the main agent. T1 supplies independently trained components, and T2 uses the fixed agent’s behavior or downstream outcomes to adapt components for that host.
5.1 T1: Agent-Agnostic Tool Adaptation
T1 supplies reusable models and specialized policies that a frozen agent orchestrates through prompts, code, or structured interfaces. Foundational Systems and Architectures illustrates this with neural operators, HuggingGPT’s model-selection workflow, ViperGPT’s Python composition, and SciToolAgent’s knowledge graph.
- HuggingGPT coordinates 1000+ models but incurs sequential-call latency and tool-description context costs. SciToolAgent accesses 500+ scientific tools through SciToolKG; these are cited system examples evaluated under different conditions.
- Categories and Training Methodologies covers vision, speech, code, retrieval, and scientific models. The authors argue that automatically evaluable outputs such as retrieval results are easier to supervise through T2 than perceptual features.
- MCP and code execution provide infrastructure for discovery and composition. Selected tool definitions and compact execution results can reduce context use; infrastructure alone does not establish that a component has been adapted through learning.
- Agents trained for query rewriting or code generation can be frozen and reused as T1 tools, linking agent training with modular deployment.
5.2 T2: Agent-Supervised Tool Adaptation
T2 adapts external components using a fixed foundation model as a source of supervision. Figure 6 and Table 3 organize methods ranging from perplexity-based retriever alignment and preference learning to multi-turn searchers, planners, and memory controllers.
- Earlier Methods: From Proxy Signals to Structured Preferences covers REPLUG’s perplexity-based alignment, AAR’s augmentation-aware preferences, and LLM-R’s distillation of an expressive preference model into a faster retriever. Proxy-tuning steers decoding through a small expert and anti-expert, while UPRISE retrieves prompts using frozen-model task performance.
- BGM places a T5-XXL bridge between a frozen retriever and PaLM2-S, using synthetic preference supervision followed by RL. The survey reports HotpotQA exact match of 35.6% versus 25.8%, a stated 38% relative improvement.
- Cascaded reformulation, retrieval, and selection separate subproblems and can defer expensive generation. Error accumulation limits deeper pipelines; the paper’s suggested 2–3-stage balance is a synthesis rather than a controlled depth comparison.
The chronology broadens T2 from retriever and interface alignment to search, planning, and memory-construction subagents. Its explicit omission of classic memory methods limits its historical coverage.
Separate tool and agent backbone columns identify the adapted model and its supervising host. This separation is necessary when interpreting systems with several LLMs and comparing total architecture with trained-component size.
Separate tool and agent backbone columns identify the adapted model and its supervising host. This separation is necessary when interpreting systems with several LLMs and comparing total architecture with trained-component size.
Separate tool and agent backbone columns identify the adapted model and its supervising host. This separation is necessary when interpreting systems with several LLMs and comparing total architecture with trained-component size.
Subagent-as-Tool extends T2 to information acquisition, memory construction, control, orchestration, and task generation. s3 trains a 7B searcher using Gain Beyond RAG, the frozen generator’s accuracy improvement over naive retrieval; the survey reports 58.9% average accuracy with 2.4k samples, roughly 70× less data and 33× faster wall-clock training than Search-R1 in the cited comparison.
- DynamicRAG learns which documents to pass and how many to retain. QAgent switches from rewarding its own answer to evaluating a stronger frozen generator’s answer, addressing retrieval of shallow, easily copied evidence.
- Mem-α trains a Qwen3-4B memory-writing controller while freezing the generator and downstream retriever, reportedly transferring from approximately 30k-token training sequences to contexts exceeding 400k tokens. AutoGraph-R1 constructs knowledge graphs using downstream reasoning utility.
- AI-SearchPlanner balances outcome, process, and cost rewards; Advisor Models supply context-specific advice, and Matryoshka Pilot learns intermediate control instructions. AgentFlow trains only a 7B planner around frozen specialists using Flow-GRPO, reporting 57.3% on search-intensive tasks, 51.5% on mathematical reasoning, and 33.1% on GAIA.
- R-Zero alternates Solver and Challenger phases, while Multi-Agent Evolve trains peripheral Proposer and Judge roles. These examples use temporarily fixed supervision and interacting learning components, requiring classification at the level of each training phase.
Agentic Memory and Skills treats memory as adaptable storage and retrieval, and skills as procedural content acquired, represented, invoked, and refined over time. Table 4 distinguishes executable code, heuristics, trajectories, API or MCP protocols, compositional programs, synthesized tools, and specialist agents by their acquisition mechanisms.
- Memory designs include hierarchical stores such as MemGPT and Memory OS, reflective stores such as Generative Agents and A-MEM, and structured stores such as HippoRAG, AriGraph, ChatDB, and Zep. CoALA and the cited forms, functions, and dynamics framework distinguish storage, content, temporal scope, and update policy.
- Memento trains a case-retrieval Q-function for a frozen GPT-4.1 planner using binary task outcomes and soft Q-learning. It reports 87.88% on GAIA validation, 79.40% on GAIA test, and 95.0% on SimpleQA, with ablations showing 4.7–9.6% absolute improvement on out-of-distribution tasks.
- MemSkill evolves memory operations; Dynamic Cheatsheet and ReasoningBank curate transferable lessons at test time without changing generator weights. Reflexion stores textual critiques, while Agent Workflow Memory retains reusable action sequences.
- Voyager retains verified executable skills and reports 63 unique Minecraft items, 3.3× the baselines. JARVIS-1 retains trajectories, EXPEL extracts heuristics, SkillWeaver stores callable APIs, and AgentStore registers specialists; LATM reports per-instance cost reductions of up to 79% through synthesized-tool reuse.
- LEGOMem extends procedural memory to teams, and ADAS searches reusable architectures. SayCan, ProgPrompt, Eureka, and RoboGen connect skills to embodied action; MemWalker, ReadAgent, CoRAG, and Adaptive-RAG address selective access to stored information.
- Memory³ and Titans illustrate parametric or hybrid boundary cases. ToolkenGPT trains tool embeddings around a frozen backbone, while UniMuR, DIFO, V2L Tokenizer, and Sysformer extend adaptation to representations, modality interfaces, or system prompts.
The rows distinguish executable functions, lessons, trajectories, protocols, programs, synthesized tools, and specialists as stored resources. Acquisition varies independently of representation, so skill format alone does not determine the learning mechanism.
6 Comparison of Adaptation Paradigms
The comparison examines cost, flexibility, data efficiency, generalization, and modularity across the four paradigms. Its quantitative evidence comes from selected published case studies rather than a shared four-paradigm experiment.
6.1 A Framework for Comparison
Table 5 compares optimization locus, supervision, cost, flexibility, and modular evolution. A1 and A2 can alter the core policy, whereas T1 and T2 can add or replace components while retaining the host’s reasoning capabilities and limitations.
- Parametric flexibility concerns changes to the agent’s policy; system-level flexibility concerns changes to external capabilities and their arrangement.
- Freezing the core prevents interference in its weights. Changes to tools or memories can nevertheless alter complete-system performance.
- The cost and generalization labels summarize tendencies in heterogeneous studies, not measured universal rankings. The survey also includes prompt-based A1/A2 adaptation, so those categories do not invariably require weight updates.
The comparison distinguishes flexibility within the core policy from flexibility through modular tools. Its cost and forgetting labels summarize architectural tendencies, with the frozen-core distinction narrower than a guarantee of unchanged system performance.
6.2 Agent Adaptation Paradigms: A1 and A2
A1 uses execution feedback to optimize tool mechanics, while A2 uses output rewards to optimize reasoning and coordination. Cited results include DeepRetrieval recall of 65.1% versus 24.7%, R1-Code-Interpreter accuracy of 72.4% on 37 test tasks, ReSearch gains of 9–22% absolute over iterative RAG baselines, and R1-Searcher improvements of up to 24% over RAG baselines.
- A1 can supply intermediate correctness information but depends on the execution environment, evaluation criteria, and reward design.
- A2 can reward search timing, evidence integration, and global coordination. Terminal rewards can also permit shortcut solutions and obscure failure sources.
- These results use different tasks, metrics, and systems and therefore do not establish a direct A1-versus-A2 ranking.
6.3 Tool Adaptation Paradigms: T1 and T2
The T1 graduation lifecycle trains a specialized agent under A1 or A2, freezes its parameters, and exposes it as a reusable tool. T2 adapts search, memory, or planning components using a fixed host’s downstream utility, potentially improving compatibility while reducing host independence.
- DeepRetrieval and SWE-Grep illustrate learned policies subsequently exposed as reusable tools.
- The proposed T2 division of labor separates information acquisition, memory construction, and planning from the host’s knowledge and reasoning.
- A component’s training paradigm and deployment role can differ. Classification needs an explicit training stage and system boundary.
6.4 Synthesis: Data Efficiency and Modularity (A2 vs. T2)
Table 6 collects selected results across domains. The central retrieval case study compares Search-R1’s roughly 170k training examples with s3’s 2.4k examples and 58.9% average accuracy; on specialized medical QA, the survey reports s3 at 76.6% versus Search-R1 at 71.8%.
- Search-R1 updates an entire Qwen2.5 agent; s3 updates a 7B searcher paired with a frozen generator that may be Qwen2.5-7B, Qwen2.5-14B, or Claude-3-Haiku.
- Optimization target, backbone composition, and architecture vary together, with training distribution another possible explanation for transfer differences. The paper explicitly states that the efficiency gap cannot be attributed solely to paradigm choice.
- The case study motivates matched experiments on narrow procedural learning around a capable frozen host. Modular replacement also requires evaluating the new component’s effect on complete-system behavior.
Retrieval recall, answer gains, memory ablations, and GAIA scores are collected under their associated signals. Different metrics prevent ranking the rows, and the s3 efficiency comparison retains the surrounding backbone and architecture caveats.
6.5 Strategic Recommendations
The authors recommend A1 for verifiable tool mechanics, A2 for integrated reasoning and coordination, T1 for reusable components, and T2 for host-aligned modular improvement. Figure 7 projects these choices onto local-to-global orchestration and monolithic-to-modular architecture, explicitly distinguishing these interpretive axes from the taxonomy’s definitions.
- A1 specialization can depend strongly on interfaces; A2 parameter updates can be expensive and affect previously learned behavior.
- T1 tools may be poorly aligned with a particular host. T2 depends on the supervising host’s quality and introduces orchestration complexity and potential error accumulation.
- Combining reusable and adaptive tools with occasional core-agent updates is the authors’ recommendation, not a demonstrated universal optimum.
This projection relates modularity to orchestration scale and depicts trained agents becoming reusable T1 tools. The authors explicitly distinguish these interpretive axes from the defining taxonomy of optimization target and supervision source.
7 Evaluation
Evaluation covers both the performance of an adapted system and the process through which its components learn. The section examines benchmark coverage, supervision signals, adaptation dynamics, and systemic requirements, then discusses missing evaluation infrastructure.
7.1 Benchmark Landscape Mapped to A1/A2/T1/T2
Table 7 maps benchmarks to task format, evaluation protocol, target, and associated paradigms. Execution assessments support A1, final-output assessments support A2, standalone tool assessments support T1, and T2 evaluation measures a tool’s marginal benefit while holding the host fixed.
- LongMemEval tests extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. The authors identify missing evaluation of skill acquisition, invocation, refinement, and drift.
- WebArena, OSWorld, τ-Bench, τ²-Bench, and other integrated suites evaluate realistic interactions, but endpoint scores alone cannot attribute gains to particular adapted components.
- The survey describes T2 evaluation as less standardized and identifies no suite supporting controlled comparisons across all four paradigms on the same task distribution.
- AAA proposes assessor agents communicating through open protocols. Its ability to enable matched comparisons at scale remains an empirical question.
The catalog connects task formats and assessment protocols with evaluation targets and paradigm labels. Entries assigned to multiple paradigms show why a benchmark name alone cannot specify whether an experiment assesses execution, final utility, or marginal tool value.
The catalog connects task formats and assessment protocols with evaluation targets and paradigm labels. Entries assigned to multiple paradigms show why a benchmark name alone cannot specify whether an experiment assesses execution, final utility, or marginal tool value.
7.2 Evaluating the Adaptation Signal
Verifiable Execution Metrics (A1 & T1) assesses tool outcomes through pass@k, test-suite pass rates, retrieval relevance, or proof checks. Holistic Utility Metrics (A2 & T2) assesses final-answer correctness, judged quality, or task completion; programmatic final-state checks can evaluate holistic outcomes.
- Execution metrics support local diagnosis but do not establish maintainability, efficiency, sound reasoning, or successful synthesis. High retrieval recall can coexist with low downstream utility, and execution metrics can be gamed.
- Formal verification provides deterministic checks within the available formal specification; it does not extend to tasks whose success criteria remain implicit or subjective.
- EM can confound answer formatting with reasoning quality. Math-Verify checks mathematical equivalence; the cited resource is https://github.com/huggingface/Math-Verify.
- LLM-as-a-judge can introduce verbosity, position, self-enhancement, and sycophancy biases. Calibration against human judgments and ensemble judging are discussed as responses.
- Final-output metrics obscure whether errors originate in reasoning, selection, execution, or synthesis, limiting attribution of a T2 tool’s contribution.
7.3 Evaluating the Adaptation Dynamics
Adaptation dynamics should be assessed through efficiency, generalization, and stability alongside final accuracy. The authors recommend reporting examples, tokens, interactions, training time, and deployment compute jointly, together with cross-task and cross-agent transfer.
- Sample and Interaction Efficiency includes interaction-count distributions and accuracy-compute Pareto frontiers. The s3-versus-Search-R1 comparison retains the confounds acknowledged in Section 6.4.
- Generalization and Robustness covers held-out tasks, different host models, noisy tools, changing interfaces, multimodal shifts, and coordination failures.
- Continual and Co-Adaptation Stability recommends retention sets and backward transfer, supplemented by policy entropy, per-update entropy changes, tool-call diversity, reward curves, and reference-policy divergence.
- Sliding-window performance variance, gradient sign changes, and convergence to equilibrium are proposed co-adaptation indicators requiring empirical validation.
7.4 Systemic Evaluation
Systemic evaluation considers cost, safety, and alignment separately from task accuracy. AgencyBench supplies a cited scale example of an average one million tokens and 90 tool calls per task, motivating separate accounting for prompts, completions, tool latency, retries, and training GPU hours.
- Cost and Inference-Time Compute recommends accuracy under token or time budgets. The suggested T2 advantage in high-latency environments remains an unvalidated cross-paradigm hypothesis.
- Safety evaluation includes sandboxed exploration, reward hacking, held-out reward distributions, trajectory inspection, and repeated safety checks during training.
- Alignment includes fidelity to user constraints and compatibility between tool outputs and host consumption. BGM’s bridge-model result illustrates agent-tool compatibility.
- Freezing a host’s weights does not prevent unsafe or misaligned behavior induced by adapted tools; safety must be assessed at the system level.
7.5 Discussion
Table 8 consolidates signal quality, dynamics, cost, safety, and alignment metrics by paradigm. The discussion explains how endpoint rankings can conceal data requirements, depleted exploration, temporary safety regressions, and improvements restricted to the chosen metric.
- Tool-centric assessment measures intrinsic component quality, while agent-centric assessment measures complete-system utility. Counterfactual evaluation varies one component while holding others fixed, but replacement can change agent behavior; progressive tool degradation is proposed as a complementary sensitivity analysis.
- The authors recommend learning curves with at least three random seeds, cost-conditioned performance, retention-set results for continual settings, and safety trajectories for RL.
- Proposed future suites compare paradigms on matched tasks and track adaptation in persistent, changing environments, including multimodal and multi-agent settings.
- Living benchmarks could maintain difficulty as capabilities improve, but introduce recurring costs, time-dependent comparability, and generated-task validation problems. The paper recommends complementing them with versioned static snapshots.
The matrix extends evaluation from signal quality to learning dynamics and systemic requirements. Primary and secondary relevance marks express recommended emphasis rather than measured coverage or the irrelevance of unmarked dimensions.
8 Applications
The applications section maps adaptation to deep research, software development, computer use, and drug discovery. Figure 8 pairs agent-side and tool-side requirements, while Table 9 relates dominant paradigms to feedback constraints and bottlenecks.
Each application pairs agent-side reasoning requirements with tool-side capabilities. The pairing explains why complete systems combine paradigms even when the survey assigns a dominant route to a domain.
8.1 Deep Research
Deep Research combines iterative search, validation, and synthesis with long-context reasoning, hypothesis refinement, and persistent knowledge tracking. The survey maps overall research quality to A2, trained search or planning components to T2, and independently trained retrievers to T1.
- Retrieval-specific measures such as Recall@K can provide A1 feedback even when complete research quality lacks a deterministic correctness test.
- The proposed domain-specialization requirements include ontologies, databases, validated computational tools, and field-specific evaluation and safety criteria.
8.2 Software Development
Software Development agents navigate repositories, modify code, debug, and validate changes across multiple stages. Tests and compilation support A1 learning, while complete issue resolution supports A2 evaluation of planning and coordination.
- SWE-Agent and OpenHands illustrate agent-computer interfaces and sandboxed development environments; SWE-Bench evaluates repository-level bug fixing.
- The survey discusses Cursor’s Tab-RL and SWE-Grep as tool-adaptation examples. Specialized retrieval can reduce irrelevant context, but SWE-Grep’s classification differs between the graduation discussion and the application section.
- Long-horizon credit assignment across multi-file edits remains a bottleneck despite execution feedback.
8.3 Computer Use
Computer Use requires grounding actions in screenshots and GUI layouts and coordinating keyboard or mouse operations. OpenCUA supplies human demonstration trajectories, while AgentTrek synthesizes trajectories from web tutorials and filters them through automatic evaluation.
- The authors map holistic task completion primarily to A2, persistent memory and context management to T2, and independently trained perception modules to T1.
- CUA-Skill stores parameterized skills and composition graphs; ACE updates structured contextual playbooks using experience.
- Intermediate GUI actions often lack the simple correctness tests available for code, although final system-state checks can verify task completion.
8.4 Drug Discovery and Development
Drug Discovery and Development combines scientific reasoning with databases, predictors, and computational utilities. Table 9 identifies T1 plug-in models as a dominant pattern and wet-lab feedback delayed by weeks to months as a constraint.
- GeneAgent uses self-verification for gene analysis, DSWizard supports inspectable analysis plans, and virtual research teams illustrate collaborative molecule design.
- Computational outputs can support A1 signals, while hypothesis generation and clinical workflow quality require broader A2 evaluation. TrialMind and LEADS illustrate evidence retrieval, TrialGPT patient-to-trial matching, and TrialGenie code execution combined with protocol evaluation.
- Biomni provides curated tools; ToolUniverse and STELLA illustrate evolving tool libraries. New scientific tools require domain-expert validation.
- The domain uses all four paradigms. Predictor scores provide computational feedback, but their validity as proxies for experimental outcomes requires separate assessment.
The profiles link adaptation choices to hard-to-verify research quality, attribution across software edits, visual grounding, and delayed wet-lab feedback. Dominant labels summarize the authors’ domain mapping and do not exclude complementary mechanisms.
9 Opportunities
The opportunities section examines joint agent-tool learning, adaptation to changing task streams, safe exploration, and efficient deployment. These are research directions derived from the survey rather than results of a unified implementation.
9.1 Co-Adaptation
Co-Adaptation asks how an agent and its tools can learn together when each changes the other’s effective environment. Figure 9 illustrates alternating A2 and T2 updates, exposing uncertain credit assignment and non-stationary optimization.
- The authors draw on co-evolutionary algorithms and multi-agent RL, including centralized training, decentralized execution, and coordination under changing partner policies.
- MATPO addresses planner-worker credit assignment within distinct prompt roles of a single LLM. Extending such mechanisms to heterogeneous learning models remains unresolved.
- Oscillation, mutual degradation, and premature convergence are possible failure modes. Relative learning-rate regulation and equilibrium analysis are proposed directions rather than validated solutions supplied by this paper.
Successive tool and agent updates use downstream output while changing different components. The illustration exposes credit-assignment and moving-target problems without reporting convergence evidence.
9.2 Continual Adaptation
Continual Adaptation distinguishes parameter updates intended to preserve past capabilities from evolving external memory and tools. Figure 10 illustrates adaptation to a changed tool; the text also considers expanding formal libraries, retrieval indices, and proof-state memories.
- Parameter-based mechanisms include regularization, orthogonal updates, low-rank adapters, routing, and model merging.
- External-memory mechanisms include replay selection, compression, tiered storage, and prompt-based retention of interaction experience.
- The cited reverse-KL and on-policy results show less forgetting than SFT in some settings. Freezing the core avoids shared-weight interference while leaving memory and tool quality to be evaluated separately.
A changed tool alters the environment in which the agent’s policy operates. The successive adaptation stages illustrate this dependency without quantifying retention of prior capabilities.
9.3 Safe Adaptation
Safe Adaptation distinguishes unsafe exploration from exploitative optimization that increases reward while undermining intended behavior. Figure 11 inserts a safety check before tool execution; the discussion covers specification gaming, adversarial tooling, and sycophancy loops.
- The cited studies report vulnerabilities in 26.1% of marketplace skills and attack success rates of up to 80% for Skill-Inject. These figures apply to those studies’ samples and evaluation conditions.
- Proposed mitigations include constrained policy optimization, safety shields, programmatic outcome checks, specification self-correction, and Proof-of-Use links between evidence and answers.
- Verifiable rewards remain dependent on adequate tests and specifications. The paper’s earlier evaluation discussion documents that execution metrics can also be gamed.
A pre-execution check rejects unsafe actions or passes accepted actions to the tool before feedback updates the policy. Its effectiveness depends on the coverage and reliability of that check.
9.4 Efficient Adaptation
Efficient Adaptation treats parameter capacity, numerical precision, and deployment location as distinct bottlenecks. Figure 12 illustrates T2 adaptation through a low-rank module; the text discusses LoRA, quantized rollouts, and local personalization.
- The cited LoRA study reports equivalence to full fine-tuning at small ranks in selected RL tasks; this does not establish equivalence for every agentic workload.
- FlashRL generates INT8 or FP8 rollouts and uses truncated importance sampling to stabilize training when rollout and training precision differ.
- Small device-specific tools could update locally around a fixed host. Combining low-rank updates, quantization, and on-device execution is presented as a composable strategy.
The low-rank module concentrates trainable capacity on the tool side while retaining a fixed host. The diagram illustrates combining T2 with parameter-efficient updates without measuring a resource or accuracy gain.
10 Conclusion
The conclusion restates adaptation as a choice of component and supervision source, connecting post-training, memory, and skills within A1/A2/T1/T2. It emphasizes signal density, multidimensional evaluation, and the reuse of trained agents as tools, while acknowledging that cross-paradigm efficiency comparisons remain uncontrolled.
- Memory and skill improvement can occur outside the core model, while parametric and hybrid mechanisms cross taxonomy boundaries.
- The authors propose stable reasoning cores with adaptive peripheral components and identify co-adaptation, continual stability, safety, and deployment efficiency as open problems.
Appendix
- The supplied PDF has no separately labeled appendix; References follow the conclusion and extend through page 89. The reviewed version is arXiv:2512.16301v3, dated 9 March 2026, with accompanying repository https://github.com/pat-jj/Awesome-Adaptation-of-Agentic-AI.
Brief Thoughts
The framework provides a practical distinction between improving execution, improving final synthesis, supplying a reusable tool, and aligning a tool to a fixed host. The paired retrieval and code examples, memory-update distinctions, and explicit confounds in the s3-versus-Search-R1 comparison provide its clearest support. Component quality and marginal downstream contribution require different assessments. The paper’s paired tool and system evaluation offers a concrete way to distinguish them. Broad claims about cost, generalization, and avoided forgetting require stronger evidence than heterogeneous case studies provide. Matched experiments are needed to test comparative predictions. Classification tensions remain: SWE-Grep appears as both a graduated T1 tool and a T2-style application example, while Self-Challenging Agents appears under A2 before receiving an A1 label in the efficiency discussion. Explicit training-stage and reward descriptions are more dependable than a label alone.