Gorio Tech Blog search

Toward Efficient Agents: Memory, Tool learning, and Planning | Summary

|

Contents

This article explains the key points of Toward Efficient Agents: Memory, Tool learning, and Planning.

  • 2026-01-20 (arXiv)
  • Yang, Xiaofang, Li, Lijun, Zhou, Heng, Zhu, Tong, Qu, Xiaoye, Fan, Yuchen, Wei, Qianshan, Ye, Rui, Kang, Li, Qin, Yiran, et al.
  • Shanghai Artificial Intelligence Laboratory, Fudan University, University of Science and Technology of China, Shanghai Jiaotong University, Institute of Automation, Chinese Academy of Sciences, The Chinese University of Hong Kong (Shenzhen), Hong Kong Polytechnic University, Wuhan University, Tsinghua University
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • Toward Efficient Agents: A Survey of Memory, Tool Use, and Planning organizes agent efficiency around three coupled components: memory, tool use, and planning. It examines how to reduce tokens, latency, computation, and interactions while preserving acceptable task quality.
  • Recurring mechanisms include compact and selectively retrieved memory, cost-aware tool invocation, adaptive deliberation, controlled search, and reusable procedures. These mechanisms can shift costs between components or from online execution to offline preparation, so their value must be assessed across complete trajectories and repeated tasks.
  • The survey evaluates efficiency through effectiveness under a fixed cost budget or cost at comparable effectiveness. It synthesizes existing methods and evaluation practices rather than introducing an algorithm or a controlled experimental comparison; inconsistent metrics and accounting boundaries limit direct comparisons between methods.

1. Introduction

Multi-step agents repeatedly retrieve context, plan, execute tools, and process observations. Because earlier outputs become later inputs, long trajectories accumulate context-processing costs, tool latency, and retries that model compression alone does not address. Figure 1 maps the surveyed literature across memory, tool use, planning, and benchmarks from 2023 to 2026.

  • The workflow is Input → [Memory → Planning → Tool Use → Observation] repeated over iterations → Solution.
  • The survey defines efficiency at the system level: preserve or improve task success while reducing token usage, inference latency, and computational cost across agent modules.
  • The categories are analytical rather than mutually exclusive. Methods are assigned according to their primary efficiency motivation, with cross-component effects discussed separately.

The four branches connect representative studies to the survey's organizational structure. The benchmark branch distinguishes effectiveness-first evaluation from efficiency-oriented signals; the chronology maps literature rather than measured performance improvement over time.

Chronological landscape of memory, tool use, planning, and benchmark research from 2023 to 2026.
Chronological landscape of memory, tool use, planning, and benchmark research from 2023 to 2026.

2. Preliminaries

The preliminaries model agent–environment interaction and distinguish agent-specific overhead from ordinary LLM generation cost. They establish why efficiency accounting must include memory, external actions, and retries across a trajectory.

2.1. Agent Formulation

An LLM-based agent is modeled as a partially observable Markov decision process augmented with tools and explicit memory: M = (S, O, A, P, R, γ; T, Ψ; Mmem, U, ρ). The environment comprises latent states, observations, actions, transitions, rewards, and a discount factor (\gamma\in[0,1)); the extensions specify tool execution and memory evolution.

  • T is the external tool set, and Ψ specifies how calls are executed and which outputs are returned.
  • Mmem is the memory state space, U maps the current memory and available information to the next memory state, and ρ is the initialization distribution.
  • This formulation is a conceptual interaction model, not a new optimization algorithm.

2.2. From Pure LLMs to Agents

Figure 2 illustrates the transition from standalone generation to trajectory-level reasoning with memory, planning, and tools. The paper approximates standalone cost as CostLLM ≈ α Ntok and agent cost as Costagent ≈ α Ntok + Itool · Costtool + Imem · Costmem + Iretry · Costretry. These expressions identify additional cost sources; they do not explicitly sum repeated operations over a trajectory.

  • Ntok counts generated reasoning tokens, and α represents per-token time or monetary cost.
  • Itool, Imem, and Iretry are binary indicators for tool invocation, memory access, and retries.
  • Agent-specific optimization therefore includes improving invocation selectivity and avoiding unnecessary recovery, alongside accelerating the underlying model.

The layout places a pure LLM at the center, surrounded by memory, tool learning, and planning within an agent system. These added capabilities introduce cost sources that faster token generation alone does not eliminate.

Extension of pure LLMs into agents through memory, tool learning, and planning.
Extension of pure LLMs into agents through memory, tool learning, and planning.

3. Efficient Memory

Memory can reduce repeated history processing, redundant exploration, and retries by preserving useful experience. Figure 3 organizes the memory lifecycle into construction, management, and access; Table 1 inventories representation-oriented mechanisms, while Table 2 covers skills and multi-agent memory.

  • Construction determines the representation and therefore influences subsequent storage, retrieval, and integration costs.
  • Management controls redundancy, staleness, and growth through updates, consolidation, deletion, and archiving.
  • Access selects memories relevant to the current decision and integrates them without recreating a long prompt.

The diagram separates representation choices, maintenance decisions, retrieval mechanisms, and integration interfaces. It identifies distinct accounting stages: compact construction may still be followed by costly semantic updates or excessive retrieval.

Memory lifecycle covering construction, management, selection, and textual or latent integration.
Memory lifecycle covering construction, management, selection, and textual or latent integration.

3.1. Memory Construction

3.1.1. Latent and Parametric Memory distinguishes activation-space storage from information encoded in a dedicated memory module’s parameters. Both reduce repeated serialization and reading of full histories, but differ in write mechanisms, capacity constraints, interpretability, and infrastructure requirements.

  • Activation Beacon compresses raw-token KV activations into beacon tokens; MemoRAG maintains window-level memory tokens; FlashMem triggers Shared-KV consolidation under high attention entropy.
  • Memory3 injects sparse explicit KV knowledge from a static knowledge base into attention. MemGen instead learns when to synthesize latent memory within the reasoning stream.
  • MemoryLLM adds a 1B-parameter layer-wise memory pool to a frozen backbone and updates it in place without backpropagation through the backbone. The survey reports that M+ extends effective retention from 20K to over 160K tokens through GPU/CPU tiers without increasing GPU memory.
  • Titans updates a neural memory module at test time through surprise-weighted gradients. Write costs are therefore method-dependent: the parametric category does not uniformly imply gradient-based updates.
  • Compact non-textual memory reduces token replay but is less interpretable and editable than text; aggressive compression can discard decision-relevant details. The M+ retention result is reported from prior work, not reproduced by this survey.

3.1.2. Textual Memory stores readable summaries, events, trajectories, profiles, plans, or symbolic facts. Table 1 separates prompt-resident, item-based, graph-based, and hierarchical approaches, whose compactness and structure can reduce online context costs while adding construction and maintenance work.

  • Prompt-resident methods maintain compact state in the active context. MemAgent overwrites summarized memory, MEM1 learns a fixed-length state marked by <IS></IS>, and AgentFold retains multi-scale summaries plus the latest full turn.
  • Item-based methods extract compact records or reusable strategies. SimpleMem resolves references into self-contained units, ReasoningBank distills successes and failures, and ACE maintains strategy bullets with usefulness statistics.
  • Graph-based methods organize entities and relations for localized or multi-hop retrieval. Zep adds temporal validity, AriGraph links semantic and episodic memory, and MAGMA uses semantic, temporal, causal, and entity graphs.
  • Hierarchical methods support coarse-to-fine access. MemGPT uses virtual-memory-style paging; MemoryOS uses storage tiers; ReadAgent links gist summaries to pages; HiAgent indexes full trajectories through subgoal summaries.
  • Text remains inspectable and editable, but retrieval and reinsertion still consume resources. More elaborate structure can improve localization while increasing indexing and update overhead.

Rows list mechanisms such as KV compression, fixed memory pools, compact textual records, graph organization, and hierarchical paging, together with method categories and resource links. The table is a mechanism inventory, not a comparison under shared experimental conditions or cost–accuracy measurements.

Memory mechanisms for latent, parametric, prompt-resident, item-based, graph-based, and hierarchical storage.
Memory mechanisms for latent, parametric, prompt-resident, item-based, graph-based, and hierarchical storage.

3.2. Memory Management

Memory management prevents an accumulating store from becoming redundant, stale, and expensive to search. The section distinguishes 3.2.1. Rule-based Management, 3.2.2. LLM-based Management, and 3.2.3. Hybrid Management by the cost and semantic adaptivity of update decisions.

  • Rule-based policies use recency, frequency, capacity, or feedback without extra LLM calls. MemoryBank uses Ebbinghaus-inspired decay; the survey cites A-MEM experiments in which forgetting-curve policies reduce memory size and retrieval time but substantially lower task performance.
  • LLM-based operation selection includes ADD, UPDATE, DELETE, and NOOP in Memory-R1 and Mem0. Memory-R1 learns these decisions through reinforcement learning, whereas Mem0 prompts an LLM after similarity retrieval.
  • Open-ended generative management, as in A-MEM, creates links and revises related notes rather than selecting only from a fixed operation set.
  • Hybrid systems use cheap triggers for routine maintenance and LLMs for semantic consolidation. MemoryOS combines tier-specific rules with LLM updates; LightMem combines online soft updates with offline sleep-time consolidation.
  • Hybrid graph management pairs semantic conflict detection with explicit graph edits. Its effectiveness depends on calibrated triggers, consistent tier policies, and reliable handling of stale facts.

3.3. Memory Access

Memory access separates 3.3.1. Memory Selection from 3.3.2. Memory Integration. Selection retrieves evidence needed for the current decision without excessive search or context growth; integration filters, compresses, formats, or injects that evidence into the agent’s computation.

  • Rule-enhanced retrieval augments semantic similarity with recency, importance, topic overlap, or attributes. RF-Mem routes familiar queries to a fast path and uncertain queries to deeper recollection; SimpleMem adapts retrieval scope to query complexity.
  • Graph retrieval anchors on query-relevant facts and expands a local subgraph, supporting entity–relation and multi-hop evidence access.
  • LLM/tool-based retrieval permits active search through memory interfaces but adds calls and latency. The survey considers it suitable for infrequent, high-stakes queries where correctness outweighs cost.
  • Hierarchical retrieval identifies a coarse segment before selectively inspecting linked details, as in H-MEM, xMemory, and HyMem.
  • Training-based retrieval learns usefulness rather than relying only on similarity. RMM learns an online reranker, Memento ranks state–case pairs with a Q-function, and MemRL combines similarity recall with Q-value selection.
  • Textual integration inserts compact working sets, strategy bullets, or adapted plan templates; RECOMP can omit unhelpful retrieval entirely. Latent integration exposes memory tokens or KV entries through internal attention, reducing text re-encoding at the expense of transparency.

3.4. Procedural Reuse via Skills

Skills are treated as procedural memory: they encode how to act in recurring situations rather than only what happened. Their efficiency benefit is amortized reuse, so future savings in planning and tool interaction must justify construction, retrieval, verification, updating, and pruning costs.

  • Trace2Skill distills trajectory-local lessons; Skill-Pro uses non-parametric policy optimization; SkillRL recursively improves a skill library and policy together.
  • AutoSkill learns from user interactions, and SKILLFOUNDRY builds procedures from heterogeneous scientific resources.
  • SkillLens selectively reuses relevant skill fragments, while Graph-of-Skills retrieves dependency-aware bundles under context constraints.
  • CoEvoSkills, SkillClaw, and SkillOS address verification and repository evolution. MemSkill extends procedural reuse to memory operations such as extraction, consolidation, and pruning.
  • An unchecked or unreliable skill library can increase context cost and error recovery. The survey therefore assesses procedural reuse across repeated tasks rather than through a single successful episode.

3.5. Multi-Agent Memory

Multi-agent memory changes the design question from what one agent remembers to where information resides, who can access it, and how updates are synchronized. Table 2 distinguishes shared, local, and mixed stores alongside procedural reuse methods.

  • Shared memory reduces duplicated exploration and history replay. G-Memory uses a three-tier graph hierarchy, RCR-Router allocates context under token budgets, and MemIndex uses intent-indexed bipartite graphs.
  • Latent sharing reduces token-level communication: LatentMAS shares latent working memory, KVComm reuses KV caches across prefixes, and Cache-to-Cache transfers source caches through a learned fuser.
  • Local memory preserves role-specific context and reduces global noise. Intrinsic Memory Agents use role-aligned templates, AgentNet prunes fixed-size stores, and DAMCS organizes local experience through a goal-oriented knowledge graph.
  • Mixed designs balance specialization and reuse. Collaborative Memory enforces provenance and sharing policies, while LEGOMem separates orchestrator-level task memory from task-agent subtask memory.
  • Centralization can introduce stale state, concurrent-write conflicts, and synchronization overhead; local isolation can cause repeated work. Mixed stores add routing and access-control complexity.

The upper block covers skill extraction, selective reuse, verification, and curation. The lower block distinguishes shared, local, and mixed multi-agent memory, showing that team-level memory design involves placement and routing as well as compact representation.

Procedural skill reuse and shared, local, or mixed multi-agent memory mechanisms.
Procedural skill reuse and shared, local, or mixed multi-agent memory mechanisms.

3.6. Discussion

The memory discussion asks whether each operation’s value for future decisions justifies its lifecycle cost. Optimizing compression, maintenance, or retrieval separately can move overhead elsewhere rather than reduce it.

  • LightMem’s cited experiments show that excessive compression reduces accuracy, whereas milder compression preserves more information at higher cost. Compression should retain decision-relevant facts, constraints, failures, and procedures rather than simply shorten summaries.
  • Online updates support immediate adaptation but increase response latency. Offline consolidation can reduce inference overhead while retaining similar total computation and delaying adaptation.
  • Evaluation should include write cost, update frequency, retrieval latency, inserted tokens, and failure-recovery savings, not recall accuracy alone.
  • Memory yields broader savings when cached plans, API constraints, successful parameter choices, or failed branches reduce subsequent tool calls and planning effort.

4. Efficient Tool Use

Tool use can shorten expensive internal reasoning, but repeated external interactions can also dominate an agent’s cost. Figure 4 divides optimization into tool selection, tool calling, and tool-integrated reasoning; Table 3 maps representative methods to these stages.

  • Selection narrows the candidate action space and reduces the tool documentation placed in context.
  • Calling addresses parameter generation, execution scheduling, search, and budget-aware orchestration.
  • Tool-integrated reasoning learns when external evidence or computation is preferable to internal knowledge.

Candidate-tool and execution-result arrows distinguish selection from calling and subsequent reasoning. Budget feedback and policy optimization show two intervention points for efficiency: controlling execution and learning more selective tool-use trajectories.

Tool-use pipeline with selection, cost-aware calling, and selective tool-integrated reasoning.
Tool-use pipeline with selection, cost-aware calling, and selective tool-integrated reasoning.

4.1. Tool Selection

Tool selection avoids placing thousands of tool descriptions in a prompt. The survey groups methods into external retrieval, multi-label classification, and vocabulary-based retrieval, with different trade-offs between low-latency access and adaptability to changing tool inventories.

  • ProTIP uses contrastive query–tool embeddings and progressively subtracts selected tool representations to avoid explicit decomposition overhead. AnyTool narrows retrieval hierarchically; DRAFT and ToolScope improve tool documentation and its use in retrieval.
  • TinyAgent uses DeBERTa-v3 small and selects tools with probabilities higher than 50%. The survey reports that inserting only selected descriptions reduces nearly half of the prompt size.
  • Tool2Vec addresses the limitations of fixed-inventory classification through two-stage retrieval and reranking. It builds embeddings from synthetic usage examples to bridge the semantic gap between queries and descriptions.
  • ToolkenGPT trains added tool embeddings while freezing other parameters; Toolken+ adds reranking and rejection; ToolGen represents each tool with a unique token.
  • External retrieval supports dynamic inventories but adds retrieval cost. Classification and vocabulary-based methods can provide low-latency selection for stable inventories, while training requirements, invocation timing, and unseen-tool generalization remain limitations.

4.2. Tool Calling

Tool calling is assessed at the trajectory level because planning or verification can prevent failed calls and retries. The section covers in-place parameter filling, parallel execution, cost-aware policies, efficient test-time search, and post-training.

  • Toolformer inserts calls into generation, while CoA uses symbolic placeholders for intermediate results. The survey reports that CoA improves performance while reducing inference time by more than 30% relative to Toolformer; this is a result from the cited study.
  • LLMCompiler schedules independent calls in parallel; LLM-Tool Compiler also fuses similar operations. W&D increases parallel information gathering within a reasoning step.
  • BTP formulates hard-budget tool allocation as a knapsack problem solved through dynamic programming. Confidence-based gating and reusable function libraries provide additional ways to reduce redundant calls.
  • ToolChain* uses A* search to prune unproductive branches; ToolTree uses MCTS-inspired search with feedback and pruning. Orchestration methods also decide when to retrieve, verify, or stop, but introduce their own control overhead.
  • OTC-PO penalizes tool use in reinforcement learning, and ToolOrchestra trains efficiency-aware orchestrators. Training-efficiency work filters low-signal prompts and down-samples less informative rollouts.

4.3. Tool-Integrated Reasoning

Tool-integrated reasoning learns the boundary between internal knowledge and necessary external execution. Supervised warm-up establishes valid tool behavior, while policy optimization improves multi-step decisions using correctness, format, parameter, and cost signals.

  • TableMind uses a plan-action-reflect loop, SFT warm-up, and Rank-Aware Policy Optimization. SMART explains whether each call is necessary, while Agent-FLAN decomposes training data into capability-specific subsets.
  • ARTIST learns from outcome-based RL; ReTool interleaves natural language with executable code; ToolRL combines format validity, correctness, and parameter matching.
  • AutoTIR and OTC-PO discourage unnecessary calls. PORTool uses step-wise importance and a decay factor γ, weighting decisions closer to the final outcome more heavily and favoring fewer tool-call steps.
  • EvoTool assigns blame across Planner, Selector, Caller, and Synthesizer modules. ELPO identifies the first irrecoverable error and converts it into localized learning signals.
  • These methods shift costs into data synthesis, sandbox execution, rollouts, and credit assignment. Poor objectives can induce tool overuse or underuse, format overfitting, or shorter but less reliable trajectories.

The inventory separates retrieval and tool-token methods from parallel execution, budget allocation, search, and RL-based reasoning. Its rows describe different intervention points rather than interchangeable solutions or ranked efficiency results.

Representative efficient tool selection, tool calling, and tool-integrated reasoning methods.
Representative efficient tool selection, tool calling, and tool-integrated reasoning methods.

4.4. Discussion

Efficient tool use means choosing external actions whose expected benefit warrants their cost, not minimizing calls indiscriminately. The discussion connects tool utility, invocation calibration, dependency-aware parallelism, and reuse through memory and planning.

  • Frequent, stable tools can use specialized low-latency selectors, while rare or changing tools remain accessible through retrieval.
  • Over-triggering wastes resources, whereas under-triggering risks hallucination or invalid action. Task quality must therefore be compared against total cost.
  • Parallel execution reduces waiting time when dependencies permit it; incorrect dependency analysis can create inconsistent results and reconciliation work.
  • Cached API constraints and successful templates reduce future selection and parameter-filling costs. Planning determines whether direct reasoning, tools, or decomposition is appropriate.

5. Efficient Planning

Planning is framed as online compute allocation: an agent decides how much reasoning, search, verification, and coordination to spend before acting. Table 4 inventories inference-time and learning-based single-agent methods alongside multi-agent coordination methods; Figure 5 connects these choices to resource constraints.

  • Single-agent efficiency controls deliberation within a trajectory and develops reusable planning capabilities.
  • Multi-agent efficiency also controls participation, communication, and parallel exploration.
  • Longer deliberation can be efficient if it prevents costly failures; a short plan can be inefficient if it repeatedly requires recovery.

Three blocks distinguish inference-time planning, learning-based policy or memory improvements, and multi-agent coordination. This organization separates methods that spend computation on the current task from those that develop reusable planning capabilities.

Planning methods grouped by single-agent inference strategies, learning-based evolution, and multi-agent collaboration.
Planning methods grouped by single-agent inference strategies, learning-based evolution, and multi-agent collaboration.

5.1. Single-Agent Planning Efficiency

Single-agent methods combine inference-time control with learning-based evolution. Figure 5 presents adaptive budgeting, structured search, decomposition, and policy or memory improvements as complementary ways to allocate deliberation.

  • Adaptive Budgeting and Control: SwiftSage separates fast behavior from slow planning. Ares selects the lowest sufficient reasoning effort per step, and Think Fast and Slow adapts cognitive depth rather than applying expensive reasoning uniformly.
  • Structured Search: LATS uses MCTS with self-reflection, CATS incorporates cost-aware pruning, and ToolChain* uses A* search. These methods guide exploration but add branching and evaluation overhead.
  • Task Decomposition: ReWOO separates planning from execution, while Task-Decoupled Planning scopes replanning to DAG-structured subgoals. Routing subtasks to specialized models can reduce unnecessary general-purpose computation.
  • Policy Optimization: QLASS uses Q-value guidance, ETO learns through DPO on trial-and-error experience, and RLTR and Planner-R1 provide process-level rewards. WebAnchor emphasizes early plan quality through dense rubric-based rewards.
  • Memory and Skill Acquisition: VOYAGER reuses learned skills, and GAP identifies parallelizable actions. Reuse can amortize planning effort but requires reliable storage, retrieval, and maintenance.

The single-agent panel connects resource constraints to budgeting, decomposition, memory, tools, and learning. The multi-agent panel depicts sparse topology, protocol/context optimization, and teacher–student distillation as separate controls over coordination overhead.

Resource-aware single-agent planning and topology, protocol, or distillation-based multi-agent efficiency.
Resource-aware single-agent planning and topology, protocol, or distillation-based multi-agent efficiency.

5.2. Multi-Agent Collaborative Efficiency

Multi-agent planning adds coordination costs: deciding who participates, what is communicated, and when interaction stops. The survey distinguishes topology optimization, protocol/context optimization, and distillation of collaboration into cheaper execution.

  • Topology optimization seeks to reduce dense message complexity from (O(N^2)) toward (O(N)) through structured or sparse interaction. Chain-of-Agents uses sequential context passing, MacNet uses DAGs, and AgentPrune removes low-utility communication edges.
  • MARS and S²-MAD limit unnecessary debate; AgentDropout and SafeSieve reduce participation or communication dynamically. InfoSeeker combines hierarchical context isolation with parallel evidence gathering.
  • Protocol optimization compresses exchanged content. CodeAgents uses pseudocode, Smurfs discards failed branches, and supervisory mechanisms terminate redundant loops.
  • MAGDI and SMAGDi distill interaction structures into a student model; D&R uses teacher–student debate to generate preference trees for DPO.
  • Aggressive pruning can remove useful disagreement or relevant context. Distillation reduces runtime coordination but incurs preparation and training costs.

5.3. Discussion

Planning efficiency is interpreted as meta-control over actions and the computation preceding them. The discussion identifies depth, breadth, and amortization as interacting axes that must be evaluated through end-to-end cost.

  • Depth controls search, reflection, and decomposition within one agent; breadth controls agents, branches, and candidate plans; amortization stores useful reasoning in policies, skills, or memories.
  • A central open problem is deciding when to stop. Fixed step counts or debate rounds do not directly estimate the marginal value of continued deliberation.
  • Memory and tools can reduce planning requirements, while planning controls their invocation. Fewer reasoning tokens do not establish lower total cost if external calls, retries, or coordination increase.

6. Benchmarks

The benchmark section adopts an effectiveness-first protocol and interprets efficiency through the effectiveness–cost Pareto frontier. Table 5 separates holistic tasks, downstream output quality, component diagnostics, and efficiency signals because existing benchmarks rarely isolate memory, tool, and planning costs completely.

  • A low-cost system that fails its tasks is not considered meaningfully efficient.
  • Comparisons should measure effectiveness under fixed cost budgets or cost at comparable task quality.
  • The section catalogs evaluation practices rather than running a common benchmark suite or reporting a new leaderboard.

6.1. Effectiveness Benchmarks

Effectiveness benchmarks establish whether agents or individual components perform their intended tasks. The survey distinguishes complete interactive trajectories from final-output evaluation and component-specific diagnostics.

  • Holistic evaluation includes GAIA, SWE-Bench, WebArena, WebShop, τ-Bench, and τ²-Bench, using task success, pass rates, correctness, and trajectory completion.
  • Downstream QA and evidence-seeking evaluation includes HotpotQA, Natural Questions, SimpleQA, BrowseComp, and SealQA. These measure final output rather than necessarily testing a complete interactive workflow.
  • Memory-specific evaluation includes LoCoMo and LongMemEval for retention, temporal reasoning, consistency, and memory-grounded answers.
  • Tool diagnostics include API-Bank, BFCL, MetaTool, ToolBench, MGToolBench, NesTools, and T-Eval. StableToolBench replaces unstable online APIs with a virtual server to improve reproducibility; MCP-RADAR and MCP-Bench add protocol-level evaluation.
  • Planning diagnostics include PlanBench, Blocksworld-style tasks, TPS-Bench, and CostBench, covering valid plans, execution, recovery, path quality, and explicit constraints.

The columns distinguish evaluation scope, representative benchmarks, and typical signals. The efficiency row overlaps with component and planning benchmarks, indicating that cost measurements generally accompany effectiveness evaluation rather than form an independent universal score.

Benchmark coverage and typical effectiveness or efficiency signals across holistic, downstream, and component-specific evaluations.
Benchmark coverage and typical effectiveness or efficiency signals across holistic, downstream, and component-specific evaluations.

6.2. Efficiency Measurements

Efficiency measurements are grouped into token and monetary cost, time, hardware resources, and interaction/search cost. Explicit accounting boundaries are necessary because reported serving latency can exclude substantial construction or maintenance work.

  • Benchmark signals include Evo-Memory’s environment-step efficiency, StoryBench’s runtime and tokens, and MemBench’s read/write times in seconds per memory operation.
  • TPS-Bench reports tokens, end-to-end time, tool-call turns, and cost-of-pass. CostBench measures Cost Gap, path deviation, and invalid tool calls.
  • Token counts approximate context-processing and API costs but do not replace hardware or runtime measurements. Cost-of-pass-style metrics connect monetary expenditure with success probability.
  • Some memory studies report search-plus-reasoning latency while excluding construction time. MemoRAG distinguishes index latency from retrieval latency, illustrating the importance of stage boundaries.
  • GPU memory, average LLM calls per response, reasoning steps, trials, explored states, and search iterations reveal overhead that token-only reporting can miss.

7. Challenges and Future Directions

The future directions focus on consistent evaluation, agentic latent reasoning, deployment-aware architectures, and multimodal agents. These are research priorities, not solutions experimentally validated by the survey.

  • Unified evaluation requires common stage definitions, metric granularity, and transparent accounting for memory, tool use, planning, latency, and runtime.
  • Agentic latent reasoning needs interfaces and objectives that accommodate tools, long-horizon memory, planning, and action verification beyond standalone text reasoning.
  • True multi-model deployments and single-model role-play pipelines should be compared under matched resource budgets because their orchestration overhead, latency, and reliability differ.
  • Multimodal agents incur perception and grounding costs in addition to language processing. Re-encoding visual histories at every step intensifies the retention–latency trade-off, while perception and grounding errors can compound across long trajectories.

8. Conclusion

The survey consolidates efficiency-oriented designs and evaluation practices across memory, tool use, and planning. It identifies recurring mechanisms and calls for standardized, transparent reporting to support fair comparison and reproducibility.

Appendix

  • The PDF contains no appendix section; the conclusion is followed by references. Public resources listed on the first page are https://efficient-agents.github.io/ and https://github.com/yxf203/Awesome-Efficient-Agents.

Brief Thoughts

The cross-component accounting perspective is particularly useful: the discussions explain how compression can create repair work, tool control can increase planning overhead, and collaboration can become dominated by communication. These examples support examining complete agent pipelines rather than interpreting local savings as end-to-end efficiency. The cited compression and forgetting trade-offs support the effectiveness-first criterion: fewer tokens or faster retrieval are insufficient when critical information is lost. The evidence supports a design and evaluation framework, not a quantitative ranking. Tables 1–4 summarize mechanisms rather than matched-budget results, and Table 5 documents heterogeneous evaluation scopes. Practical comparisons should distinguish online latency from total computation and assess reusable skills or caches across enough repeated tasks to account for construction and maintenance.