Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments | Summary
12 Feb 2026 | Paper Review LLM Evaluation Synthetic Environments Multi-Agent CoordinationContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 ARE: Scaling Up AGENT Environments and Evaluations
- 4 GAIA2: Expanding General AGENT Evaluation
- 5 EXPERIMENTS
- 6 Conclusion & Discussion
- Appendix
- Brief Thoughts
This article explains the key points of Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments.
- 2026-02-12 (arXiv), ICLR 2026
- Froger, Romain, Andrews, Pierre, Bettini, Matteo, Budhiraja, Amar, Cabral, Ricardo Silveira, Do, Virginie, Garreau, Emilien, Gaya, Jean-Baptiste, Laurençon, Hugo, Lecanu, Maxime, et al.
- Meta SuperIntelligence Labs
- Paper
Summary
- Gaia2 evaluates LLM agents in simulations that evolve independently of their actions, including during model inference. Built on Agents Research Environments (ARE), it contains 800 human-authored scenarios across five core capabilities and 320 augmented scenarios testing noise robustness and Agent2Agent collaboration.
- The ARE Verifier checks required write actions for argument consistency, causal ordering, timing, and completeness while leaving exploratory reads outside oracle matching. On 450 hand-labeled trajectories, it achieves 0.98 agreement, 0.99 precision, and 0.95 recall, supporting evaluation and reinforcement learning from verifiable rewards (RLVR).
- Under the shared ReAct scaffold, GPT-5 (high) achieves the highest overall pass@1 at 42.1%, but scores 0.0% on Time; removing generation latency raises its Time score to 34.4%. Collaboration improves token-normalized repeated-sampling performance for Llama 4 Maverick but not Claude 4 Sonnet, demonstrating model-dependent trade-offs among accuracy, latency, and coordination.
The curves count successful recorded runs below each cost threshold, so model rankings depend on the allowed budget. GPT-5 (high) reaches the highest plateau, while cheaper configurations are competitive at tighter thresholds. The plateaus reflect remaining failures in the recorded evaluations, not an experiment that grants each run progressively more inference compute.
1 Introduction
The paper addresses the mismatch between synchronous benchmarks and deployed agents that must handle delayed messages, changing conditions, deadlines, and other agents. Its three contributions are the ARE simulation framework, the Gaia2 benchmark, and an empirical comparison of proprietary and open-source models that measures accuracy alongside cost and execution time.
- Gaia2 generalizes information seeking beyond web-only tasks to a controlled smartphone-like environment with email, messaging, calendars, contacts, and other apps.
- Its write-action annotations expose intermediate mistakes that final-answer or final-state evaluation can overlook.
2 Related Work
The related-work discussion examines both the capabilities covered by agent benchmarks and their verification methods. Gaia2 combines asynchronous environment dynamics with action-level verification.
Benchmarking LLM agents
The paper situates Gaia2 alongside grounded execution benchmarks such as ALFWorld, WebShop, WebArena, and WorkArena; app-based environments such as AppWorld and ToolSandbox; and function-calling, temporal, and multi-agent evaluations. Gaia2 lets external events occur independently of agent actions and evaluates timing, ambiguity, robustness, and coordination within one controlled environment.
- GAIA, SWE-bench, and BrowseComp are discussed as evaluations emphasizing final outcomes.
- The comparison motivates broader coverage, but does not establish that every cited benchmark lacks all forms of intermediate verification or temporal interaction.
Verification in agentic benchmarks
GAIA uses final-output exact matching, whereas ToolSandbox constrains trajectories through milestones and minefields; rubric-based approaches support judgments beyond strictly exact answers. The ARE Verifier combines exact argument checks, tool-specific LLM judgments, and causal and temporal constraints to assess every required write action.
- This hybrid design allows flexibility in message content without delegating the entire trajectory judgment to an LLM.
- The verifier is presented as reusable infrastructure for other ARE environments, not only as Gaia2-specific scoring code.
3 ARE: Scaling Up AGENT Environments and Evaluations
ARE decouples the environment loop from the agent loop: simulation time advances and scheduled events execute while an agent is reasoning. The platform records interactions and state changes so that researchers can reproduce scenarios, inspect trajectories, and analyze failures.
Core concepts
ARE defines five interacting abstractions: apps, environments, events, notifications, and scenarios. A scenario specifies an initial state and an event DAG containing requests, external changes, and verification logic; independent branches need not follow a total ordering.
- Apps expose stateful APIs whose tools are classified as read-only or write; environments combine apps with time management and governing rules.
- Events are timestamped and logged, with absolute or relative scheduling and explicit dependencies. Notification policies control which events become visible to the agent.
- Verification can run online at milestones or offline after execution. The authors also report reimplementing τ-bench, τ²-bench, GAIA, BFCL-v3, and VendingBench within ARE.
The event queue and event loop operate separately from user and agent interaction. Logged events feed notification policies and verification, connecting asynchronous observation with auditable scoring.
Asynchronicity and time
Model generation consumes simulated time, so a slow response can miss a deadline or arrive after the environment has changed. Responsiveness therefore contributes directly to task success.
Mobile environment
Mobile instantiates ARE with twelve apps and 101 tools. Each user-centered universe contains 400K–800K tokens of structured and unstructured content, excluding filesystem contents, generated from PersonaHub-seeded personas and cross-app dependencies.
- A turn begins with a user message or notified event and ends when the agent replies to the user; scenarios also terminate on completion, limits, or verification failure.
- Figure 3 shows broad app usage for Llama 4 Maverick, with Contacts accounting for 19.0%, Messages for 15.8%, and Emails for 11.8%.
- Mobile is API-based, not an evaluation of touchscreen or visual interface control. The appendix identifies unresolved cross-app consistency issues despite the dependency-based generation process.
Llama 4 Maverick interacts with all twelve Mobile apps rather than concentrating on a single service. Contacts has the largest share at 19.0%, followed by Messages at 15.8%, consistent with tasks requiring cross-app information gathering.
Agent orchestration
The baseline orchestration uses a model-agnostic ReAct loop that emits one structured JSON tool call per step. Pre-step hooks inject queued notifications into context, and post-step hooks check termination, allowing interaction with asynchronous and multi-turn scenarios.
- A Parallel Tool Calling ablation improves efficiency in some configurations but does not consistently improve pass@1.
- These results weaken the explanation that sequential tool calling alone causes the low scores, but do not establish that model capabilities are the only bottleneck or rule out better orchestration.
4 GAIA2: Expanding General AGENT Evaluation
Gaia2 contains 800 unique, human-authored scenarios across 10 Mobile universes, with 160 scenarios for each core capability. Gaia2-mini is a 160-scenario subset; applying Noise and Agent2Agent augmentations to that subset adds 320 scenarios, yielding 1,120 in total.
4.1 Capabilities Evaluated
The five core capabilities are Execution, Search, Ambiguity, Adaptability, and Time. These are dominant task flavors rather than orthogonal dimensions: a timed task can also require substantial search and execution.
- Execution emphasizes sequences of write actions; Search emphasizes gathering information across apps; Ambiguity tests reporting impossible, contradictory, or underspecified requests.
- Adaptability requires revising behavior after environmental changes, while Time requires acting within specified temporal constraints. Artificial combinations of three or more capabilities were excluded because early experiments produced unnatural tasks with unclear evaluation signals.
- Noise introduces tool failures, signature changes, and distractor events. Agent2Agent replaces direct app access with on-demand app-agents, testing decomposition and coordination under partial observability; the default main-agent and app-agents use the same model.
The examples distinguish the five core task flavors from two environment augmentations. Noise modifies execution conditions, while Agent2Agent changes how apps are accessed; both can reuse existing task annotations.
4.2 Scenario Design and ANNOTATION Protocol
Annotators explore generated universes and construct ground-truth DAGs of required write actions and environment events using the ARE annotation interface. Independent validation rounds, consistency checks, structural guardrails, and baseline-agent difficulty calibration are used to reduce annotation errors.
- Tasks target a dominant capability while retaining natural combinations of reading, writing, and environmental interaction.
- The oracle specifies required write behavior rather than a prescribed sequence of exploratory reads.
4.3 Verifier
The verifier first requires matching write-tool names and counts, then maps oracle actions to agent actions while checking arguments, dependencies, and timing. Success requires matching every oracle write action; independent goals may be completed in either order.
- Rigid fields such as IDs use exact matching, while flexible fields such as email text use tool-specific LLM rubrics and anti-hacking checks.
- Table 1 reports 0.98 agreement, 0.99 precision, and 0.95 recall for ARE, compared with 0.72, 0.53, and 0.83 for an LLM-only in-context verifier on the same 450 trajectories.
- Read actions are excluded from oracle matching, but remain subject to the experiment’s step, context, and time limits.
The structured verifier raises precision from 0.53 to 0.99 and agreement from 0.72 to 0.98 relative to an LLM-only trajectory judge. Recall also rises from 0.83 to 0.95, indicating that the precision gain is not achieved simply by rejecting nearly all trajectories.
5 EXPERIMENTS
The experiments compare models across all seven splits and separately investigate Time latency, Agent2Agent collaboration, and selected orchestration settings. Overall pass@1 is the arithmetic average across splits, not a score restricted to the five core capabilities.
Experimental setup
The main setup uses full model context lengths of at least 128K tokens, a stated temperature of 0.5, 16K token generation limits per turn, and three runs per scenario. Runs stop at 200 steps, context overflow, verifier completion or failure, or timeout; the default notification policy is medium.
- Appendix B.4 specifies an exception for GPT-5: temperature and top-p are set to 1 because of API constraints.
- The evaluation pauses during responses and resumes with matching simulated generation-time offsets to accommodate deployment issues while retaining generation latency in the simulation.
- The verifier uses Llama-3.3-70B-Instruct at temperature 0. Intermediate reasoning outputs are discarded between steps, which the authors acknowledge may be suboptimal for models that interleave reasoning and tool use.
5.1 Core Results
Table 2 gives GPT-5 (high) the highest overall score, 42.1%, followed by Claude-4-Sonnet Thinking at 37.8% and Claude-4-Sonnet at 34.8%. Kimi-K2 leads the evaluated open-source models at 20.1%; the abstract and some discussion passages instead describe its score as approximately 21%.
- GPT-5 (high) scores 69.2% on Execution, 79.6% on Search, 51.9% on Ambiguity, and 35.4% on Noise, but 0.0% on Time and 17.9% on Agent2Agent.
- Claude-4-Sonnet Thinking scores 42.1% on Adaptability, 8.5% on Time, and 32.5% on Agent2Agent, illustrating the capability-specific rankings in Figure 5.
- Execution and Search are relatively easier in this benchmark, but their strongest scores do not establish complete task saturation.
The overall leader is not the leader in every split. GPT-5 (high) achieves 42.1% overall but 0.0% on Time, whereas Claude-4-Sonnet Thinking reaches 8.5% on Time and 32.5% on Agent2Agent; Kimi-K2's exact overall score is 20.1%.
Independent reranking makes changes in model ordering visible across capabilities. The low Time bars contrast with stronger Execution and Search scores, while Claude's Thinking configuration leads the displayed Adaptability and Agent2Agent comparisons.
Increasing GPT-5 reasoning effort improves overall accuracy while increasing solution time. Claude 4 Sonnet costs roughly 3× more than GPT-5 (low) for comparable accuracy but runs substantially faster; pricing estimates use Artificial Analysis data accessed September 10, 2025.
- Figure 1 counts successful recorded runs whose cost falls below each budget threshold. It does not test assigning progressively larger inference budgets to the same model.
- Figure 6 places Kimi-K2 among the cost-effective models and Grok-4 among the less efficient configurations.
- Human annotators are reported to solve every task but take longer through the ARE GUI, so the human timing comparison also reflects interface differences.
Accuracy, monetary cost, and completion time produce different model orderings. Claude 4 Sonnet offers shorter successful-scenario times at higher cost than GPT-5 (low), while the human timing reference reflects interaction through the ARE GUI rather than a matched model interface.
Across evaluated models, pass@1 correlates positively with average model calls and generated output tokens, as shown in Figure 7. Claude-4 Sonnet and Kimi-K2 achieve comparatively strong scores with relatively few output tokens, while Thinking variants of Claude and Qwen use more tokens per step but fewer steps overall.
- These are cross-model associations, not controlled evidence that adding tool calls or tokens by itself causes higher success.
- The authors report nearly identical app usage patterns across models; their proposed architectural explanations for token efficiency are not tested directly.
Higher scores generally accompany more model calls and output tokens, but Claude 4 Sonnet and Kimi-K2 attain relatively high accuracy with fewer generated tokens. These scatterplots identify associations and efficiency outliers, not causal effects of increasing compute.
5.2 TIME REVEALS THE IMPACT OF INFERENCE SPEED—AND SYSTEM RELIABILITY
Removing generation latency in instant mode raises GPT-5 (high) from 0.0% to 34.4% on Time, Claude-4 Sonnet from 8.1% to 26.7%, and Gemini 2.5 Pro from 7.3% to 14.8%. In default mode, higher GPT-5 reasoning effort improves Execution performance while reducing temporal responsiveness.
- The latency ablation shows that generation time contributes substantially to Time failures under the tested setup, although instant-mode scores remain far below complete success.
- Adaptive compute and reliable serving infrastructure are proposed directions; neither an adaptive policy nor a controlled serving-reliability intervention is evaluated.
- Some tasks require concurrent actions within narrow windows that the baseline single-threaded scaffold cannot fully express.
Removing generation latency substantially improves Time performance, particularly for GPT-5 (high), whose score rises from 0.0% to 34.4%. The GPT-family comparison also shows that better Execution scores can coexist with worse deadline-sensitive performance.
5.3 A Closer Look at Multi-AGENT Collaboration on GAIA2 with Agent2Agent
On Gaia2-mini, increasing the Agent2Agent ratio r improves Llama 4 Maverick’s repeated-sampling pass@k versus total-token trade-off and reduces tool-call errors. Claude 4 Sonnet does not gain a comparable token-normalized advantage, and its tool-call error rate is higher in the collaboration settings than in the single-agent setting.
- The main-agent delegates temporally extended subtasks to app-agents, introducing hierarchical decomposition and message-passing overhead.
- Figures 9 and 10 support model-dependent collaboration benefits, not a general claim that more collaborators improve every agent.
- The proposed explanation—that delegation helps when decomposition benefits exceed coordination costs—is not isolated through a dedicated causal experiment.
The delegation example shows a main-agent specifying a subtask and receiving an app-agent report. Increasing delegation reduces tool-call errors for Llama 4 Maverick. Claude 4 Sonnet instead has higher error rates in both collaboration settings than in its single-agent setting, with no monotonic improvement across ratios.
At comparable total-token expenditure, collaboration improves Llama 4 Maverick's repeated-sampling success curve. Claude 4 Sonnet shows no comparable token-normalized advantage, limiting the result to particular model and orchestration combinations.
Table 3 evaluates heterogeneous main-agent/app-agent pairings at r = 1, averaged over three runs with standard errors. Its labeled matrix gives all-Llama 8.5 ±1.7, Llama-main with Claude-app 18.3 ±0.7, Claude-main with Llama-app 16.2 ±0.7, and all-Claude 29.3 ±2.9.
- Replacing either role with Claude improves over the all-Llama configuration; the strongest tested team uses Claude in both roles.
- The surrounding prose swaps the two mixed-team scores. The values here follow the table’s main-agent columns and app-agent rows.
- These results do not demonstrate that a heterogeneous team outperforms the strongest homogeneous team or is optimal after monetary cost normalization.
The labeled matrix assigns 18.3 ±0.7 to Llama-main/Claude-app and 16.2 ±0.7 to Claude-main/Llama-app, despite the surrounding prose swapping these values. Both mixed configurations exceed all-Llama at 8.5 ±1.7, while all-Claude remains strongest at 29.3 ±2.9.
6 Conclusion & Discussion
The paper concludes that asynchronous evaluation exposes accuracy–cost–latency trade-offs hidden by static task scores. ARE supplies reusable action-level verification for evaluation and RLVR, while the Time and Agent2Agent experiments show the importance of evaluating responsiveness and delegation alongside task accuracy.
- The verifier’s measured precision and recall support its practical use, but the observed reward-hacking episode shows why validation and adversarial checks remain necessary.
- Adaptive compute, hybrid reward signals, and improved delegation are proposed directions rather than demonstrated solutions.
Appendix
- A ARE APPENDIX and A.1 ARE FOUNDATIONS describe apps, environments, events, notifications, and scenarios. Apps maintain internal state; Python App methods become tool descriptions, decorators distinguish read and write tools, and role scoping separates agent, user, and environment actions. MCP compatibility and replaceable storage backends support extension. AgentUserInterface offers blocking and non-blocking communication, while System supplies time queries and waiting tools. Waiting accelerates simulation to event-to-event execution. A.1.2 ENVIRONMENT defines a Markov Decision Process whose state includes apps, time management, and notifications; execution is deterministic given a fixed initial state and seed.
- A.1.3 EVENTS describes event creation, dependency-aware scheduling in EventQueue, execution in EventLoop, and recording in EventLog. Tool-call events represent agent, user, or environment activity; conditional events monitor criteria, validation events enforce milestones or constraints, and oracle events encode ground truth. A.1.4 NOTIFICATION SYSTEM and A.2 NOTIFICATION POLICIES IN ARE vary observability through low, medium, and high policies, while user messages are always notified; agents can also proactively inspect app contents without an alert. The medium policy includes selected email, messaging, shopping, ride, and calendar changes. A.1.5 SCENARIOS implements initial states, scheduled events, and verification in scenario.py, typically beginning with send_message_to_agent. Its two-turn example switches from searching Chats to processing an asynchronously received email, with validation at turn boundaries.
- A.3 UNIVERSE GENERATION gives foundational apps such as Contacts priority when resolving dependencies. PersonaHub seeds are expanded with location information, and a sampled user persona anchors each universe. Chats and Messages include group and individual conversations with at least 10 alternating messages; Emails generates inbox and sent content from user and contact personas. Calendar iteratively generates weekly events and summaries, RentAFlat supplies listings, City assigns zip-code crime values in the range 1–100, and Cabs generates country-conditioned ride histories. Shopping samples 500 products per universe from a public Amazon dataset; Files shares public content and files without personal information across universes. Unresolved temporal consistency, relationship consistency across communication channels, and cross-modal references limit synthetic-world coherence.
- A.4 ARE GRAPHICAL USER INTERFACE combines environment exploration, event-DAG visualization, timestamped trace inspection, and replay after editing a past step. App views are generated automatically; scenario views expose dependencies, triggers, timing constraints, and execution status, with time jumps for long scenarios. Unlike trace-only annotation tools, the interface supports scenario construction within a reproducible simulation. Its graph-based annotation editor checks structural and timing inconsistencies and is reported to improve Mobile annotation time by approximately five times relative to manual approaches. The annotation interface is explicitly not released at the time of the paper.
- B.1 DETAILS OF GAIA2 ANNOTATION requires connected event DAGs rooted at send_message_to_agent, a single conversation branch, and turns ending with send_message_to_user; only user-message or environment events may follow that turn-ending action. Examples cover updating contacts aged 24 or younger, aggregating friends’ cities across Contacts and Chats with alphabetical tie-breaking, detecting scheduling conflicts, adapting a property visit to a friend’s reply, and ordering a cab after 3 minutes without a response. Search permits only the final answer as a write action and requires concise answers without complex computation. Ambiguity is single-turn, explicitly requests reporting ambiguities, and includes impossible requests, contradictions, missing information, and multiple valid interpretations. Adaptability follows an act–report–dependent-environment-event–adapt pattern; distractors can be added. Event taxonomies distinguish independent events rooted in the initial request from dependent events following the agent’s report, with structural exceptions for distractors. Time covers one-off and recurrent tasks triggered by clocks or events, within a five-minute horizon. Timing checks apply when the oracle relative delay exceeds 1 second, accepting actions within ([\Delta t-5\,\mathrm{sec},\;\Delta t+25\,\mathrm{sec}]); actions intended immediately after an event are annotated with a +2 sec delay.
- B.2 VERIFICATION DETAILS checks write-tool counters, topologically orders oracle actions, and matches arguments, preceding parent actions, and relative timing. It assumes no equivalent alternative write actions: using Messages instead of the oracle’s Chats tool fails. Multi-turn verification accepts a trajectory only if every turn passes, either online or offline; execution can gate later turns on successful verification, while test scenarios without accessible oracles advance after send_message_to_user. RL experiments exposed judge hacking through code-like answer content, and an added task-agnostic style check addressed the observed exploit. Verifier validation combines validity-preserving and validity-breaking oracle perturbations with 450 manually labeled model trajectories, since synthetic unit tests cannot cover every real behavior. Llama 3.3 70B Instruct powers both the structured verifier and the in-context baseline. Alternative verifier models Gemini 2.5 pro and Claude Sonnet 3.7 each achieve 0.96 agreement, 0.98 precision, and 0.89 recall on the same validation set.
- B.3 AGENT ORCHESTRATION defines each ReAct step through Thought, Action, and Observation, surrounded by configurable pre-step and post-step methods. Parallel Tool Calling changes pass@1 by −6.3 to +3.0 percentage points in the reported Execution and Time ablation; GPT-5 (low) on Execution saves 435 s and 5109 output tokens but loses 1.0 percentage point, while some configurations consume more time or tokens. B.4 EXPERIMENTAL SETUP AND IMPLEMENTATION DETAILS uses custom stop sequences for Claude, Kimi, and Qwen, enables Gemini dynamic reasoning, caps Grok-4 reasoning at 16K tokens, and sets GPT-5 temperature and top-p to 1. Frequent Grok-4 Empty Response errors, model exclusions due to cost and latency, and discarded intermediate reasoning constrain how closely these scores approximate each model’s best agent performance.
- B.5 ADDITIONAL EXPERIMENTS reports broadly similar sub-agent spawning behavior across model families, with stronger Agent2Agent performers tending to instantiate more collaborators; no explicit collaborator-count limit applies before scenario timeout. Its noise ablation varies tool-error probabilities and distractor-event frequencies, giving Claude-4 Sonnet scores of 31.2 without noise, 35.0 with low noise, 23.8 with default medium noise, and 8.1 with high noise. The displayed table covers one model and provides no uncertainty estimates, so it supports degradation at stronger noise levels for that model rather than a quantified universal trend or a statistically established low-noise benefit.
Brief Thoughts
The latency ablation is particularly informative: the overall leader fails every default Time scenario yet recovers substantially when generation time is removed. Gaia2 separates task accuracy from timely execution, although its synthetic API environment, structured prompts, and five-minute Time horizon limit claims about general deployment reliability. Action-level verification detects incorrect intermediate writes, but exact write-tool counts and tool identity also reject alternative or corrective strategies that could achieve a user’s broader goal. The collaboration results support reporting model roles and resource-normalized outcomes rather than treating agent count as a universal scaling variable. RLVR compatibility is an infrastructure contribution. The reported RL experiments reveal verifier hacking, but do not show a trained agent closing the benchmark’s performance gap.