Gorio Tech Blog search

Toward Autonomous Long-Horizon Engineering for ML Research | Summary

|

Contents

This article explains the key points of Toward Autonomous Long-Horizon Engineering for ML Research.

  • 2026-04-14 (arXiv)
  • Chen, Guoxin, Chen, Jie, Chen, Lei, Zhao, Jiale, Meng, Fanzhe, Zhao, Wayne Xin, Song, Ruihua, Chen, Cheng, Wen, Ji-Rong, Jia, Kai.
  • Gaoling School of Artificial Intelligence, Renmin University of China, Independent Researcher, AweAI Team
  • Paper
  • Github

Read this article in Korean


Summary

  • Toward Autonomous Long-Horizon Engineering for ML Research studies how agents convert underspecified research objectives into runnable, experimentally validated ML systems. The central problem is preserving decision-relevant project evidence across transient agent invocations so that implementation, experimentation, and refinement remain cumulative.
  • AiScientist combines hierarchical specialist delegation with a schema-governed File-as-Bus workspace. With a 24-hour budget and one H20 GPU per task, it achieves PaperBench average scores of 30.52 with Gemini-3-Flash and 33.73 with GLM-5, and 81.82 ±0.00 Any Medal% on MLE-Bench Lite with both backbones.
  • Under GLM-5, removing File-as-Bus reduces PaperBench score by 6.41 points and MLE-Bench Lite Any Medal% by 31.82 percentage points while preserving submission validity. These ablations support durable state as a contributor to refinement, but repeated PaperBench evaluation is constrained by grading cost, and comparisons with Codex/GPT-5.5 xhigh do not isolate architecture from backbone differences.

1. Introduction

The introduction distinguishes research engineering from isolated idea generation, code generation, or scientific writing: implementations must survive environment setup, resource integration, delayed experiments, and diagnosis of confounded failures. Figure 1 illustrates this iterative setting on Detecting Insults, where AiScientist completes 74 experiment cycles over 23 hours without human intervention and raises validation AUC from 0.903 to 0.982 through 18 best-so-far updates.

  • Role continuity means that a specialist invoked again can recover earlier assumptions, attempted fixes, failures, and unresolved questions.
  • Project continuity means that different specialists coordinate through shared evidence rather than compressed conversational handoffs.
  • The example demonstrates sustained validation improvement on one task; it is not an aggregate estimate of generalization or benchmark performance.

The running-best curve advances through intermittent successful experiments rather than improvement in every candidate. Its 18 retained updates among 74 cycles illustrate repeated refinement interspersed with unsuccessful trials; the reported AUC is a validation metric on one task.

Autonomous refinement on Detecting Insults over 23 hours, raising validation AUC from 0.903 to 0.982 across 74 experiment cycles.
Autonomous refinement on Detecting Insults over 23 hours, raising validation AUC from 0.903 to 0.982 across 74 experiment cycles.

2. Task Formulation

The task takes a research specification, execution environment, resource-access policy, and time budget as inputs, and requires a submission evaluated by fresh execution in a clean environment. Success includes both executability and empirical behavior consistent with the target objective.

  • Underspecification requires recovering missing implementation decisions from the specification, related literature, and permitted public resources.
  • System Setup Burden includes dependency configuration, dataset and model acquisition, and integration into a runnable pipeline.
  • Delayed Feedback complicates diagnosis because discrepancies may originate in interpretation, implementation, preprocessing, or infrastructure.
  • State Continuity requires later decisions to correctly reuse code, configurations, logs, results, and diagnostic evidence.

3. AiScientist

AiScientist couples a durable artifact protocol with a hierarchical research team. Detailed project evidence remains in the workspace, while the Orchestrator routes work through compact directives and summaries.

3.1. Overview: Thin Control over Thick State

Thin control over thick state separates the Orchestrator’s active control context from the project’s persistent record. Figure 2 shows the Orchestrator issuing stage-level directives and receiving concise summaries, while specialists use a compact workspace map to locate and update relevant artifacts.

  • File-as-Bus preserves detailed evidence but does not itself determine which bottleneck should be addressed next.
  • Hierarchical orchestration selects the next role and intervention without loading the complete project history into every control decision.
  • The two mechanisms jointly target role continuity and project continuity across repeated specialist invocations.

The architecture separates a compact control channel from detailed persistent state. Specialist directives and summaries pass through the Orchestrator, while analyses, plans, runnable code, and append-only logs support subsequent invocations and cross-role coordination.

AiScientist’s hierarchical research team and File-as-Bus workspace, with specialist roles, scoped subagents, tool access, and persistent artifacts.
AiScientist’s hierarchical research team and File-as-Bus workspace, with specialist roles, scoped subagents, tool access, and persistent artifacts.

3.2. File-as-Bus Coordination

File-as-Bus is an artifact-mediated coordination protocol, not simply a shared directory. With durable workspace state W_t and artifact schema σ, the workspace map is defined as (m_t = \mathcal{M}(W_t; \sigma)), a compact index combining available artifacts with their purposes and access contracts.

  • Append-only logs retain chronological implementation decisions, experiment outcomes, failures, and diagnoses.
  • Versioned artifacts are specified to expose a canonical current state with prior revisions, while mutable state covers information that does not require the same history; Appendix B describes revision preservation conceptually rather than detailing a version-control mechanism.
  • Table 3 specifies artifact purposes, update modes, writers, and readers. Runnable repository files are mutable, with their decision histories recorded in logs rather than by versioning every file.

3.3. Hierarchical Research Team

The Orchestrator treats delegation as Agent-as-Tool: (a_t = \pi_0(c_t, m_t)), with (a_t \in \mathcal{T}0 \cup \mathcal{A}_1), allowing either native tool use or invocation of a Tier-1 specialist. A specialist returns a concise summary and durable updates through ((s_t, \Delta W_t) = \pi_j(d_t, m_t; W_t)), followed by (W{t+1} = \mathcal{U}(W_t, \Delta W_t)).

  • Paper Comprehension extracts methods, target metrics, baselines, ambiguities, and implementation assumptions; Prioritization orders work by dependencies, impact, and feasibility.
  • Implementation builds or patches the system and integrates resources; Experimentation executes it, compares results with objectives, and records diagnoses.
  • A Generic Helper Interface supports bounded auxiliary work. Tier-2 subagents do not recursively delegate; their findings return to the invoking specialist and become durable artifacts when relevant.

3.4. Evidence-Driven Research-Engineering Loop

The loop first constructs an execution plan and runnable scaffold, including code, configuration, setup, resource acquisition, and an entry point. It then alternates implementation and experimentation, using recorded failures, partial successes, metric gaps, and discrepancies to select the next intervention.

  • Possible interventions include code patches, preprocessing or configuration changes, reruns, and renewed paper analysis when an assumption is invalidated.
  • Failures become persistent diagnostic evidence, while successful changes remain available as runnable artifacts and result records.

4. Experiments

The experiments evaluate from-scratch research-paper replication on PaperBench and competition-style iterative ML improvement on MLE-Bench Lite. Time-dependent trajectories, ablations, and workflow analyses supplement final outcomes to examine sustained refinement beyond initial executability.

4.1. Experimental Setup

Each task receives one H20 GPU and a 24-hour budget, with AiScientist instantiated using Gemini-3-Flash or GLM-5. PaperBench contains 20 replication tasks and uses the official full-evaluation protocol with GPT-5.4 grading; MLE-Bench Lite contains 22 tasks and reports mean ± SEM over three runs/seeds.

  • PaperBench forbids the authors’ original code and other blacklisted resources; permitted external resources remain available for constructing a fresh implementation.
  • Matched PaperBench baselines are BasicAgent and IterativeAgent. Controlled MLE-Bench Lite comparisons use AIDE, LoongFlow, and ML-Master 2.0 in the reported backbone-specific configurations.
  • Official leaderboard entries and Codex/GPT-5.5 xhigh are contextual references, not matched backbone comparisons.
  • A full 20-task PaperBench grading pass costs approximately $832, materially constraining repeated full-benchmark evaluation.

4.2. Main Results on PaperBench

Table 1 reports PaperBench average scores of 30.52 for Gemini-3-Flash and 33.73 for GLM-5, exceeding the strongest matched baselines by 9.92 and 11.15 points, respectively. AiScientist exceeds both matched baselines on every listed task within each backbone block, although absolute scores and improvement magnitudes vary substantially.

  • Under Gemini-3-Flash, average cost per task is $15.67 for AiScientist versus $27.44 for IterativeAgent; under GLM-5, the corresponding costs are $12.20 and $54.90.
  • AiScientist is not the cheapest system: BasicAgent costs $6.25 and $4.90 per task under the two backbones, respectively.
  • The GLM-5 result exceeds the Codex/GPT-5.5 xhigh reference score of 29.45 by 4.28 points. This is a cross-system comparison rather than an isolated architectural effect.
  • IterativeAgent’s higher cost but lower score argues against interaction volume alone explaining the gains; it does not independently establish which component causes them.

Every listed task shows a positive gain over the stronger matched baseline under both backbones. Average scores reach 30.52 and 33.73, while costs are lower than IterativeAgent but higher than BasicAgent, showing a performance–cost trade-off rather than uniformly minimal cost.

Full PaperBench results across 20 tasks, including backbone-matched baselines, average scores, per-task costs, and a Codex/GPT-5.5 xhigh reference.
Full PaperBench results across 20 tasks, including backbone-matched baselines, average scores, per-task costs, and a Codex/GPT-5.5 xhigh reference.

4.3. Main Results on MLE-Bench Lite

Table 2 reports 100.00 ±0.00 Valid Submission% and 81.82 ±0.00 Any Medal% for AiScientist under both backbones. Any Medal% exceeds LoongFlow with Gemini-3-Flash by 4.55 percentage points and ML-Master 2.0 with GLM-5 by 16.67 percentage points.

  • Above Median% is 86.36 ±0.00 with Gemini-3-Flash and 89.39 ±1.52 with GLM-5, improving over the strongest corresponding controlled baselines by 9.09 percentage points.
  • Codex/GPT-5.5 xhigh obtains 68.18 ±2.62 Any Medal%, 13.64 percentage points below AiScientist.
  • AiScientist exceeds the highest official leaderboard Any Medal% shown, 77.27 ±0.00 for AIBuildAI, but these external rows use different backbones and reporting contexts.
  • Any Medal% is the primary endpoint. AiScientist does not dominate every medal tier, so the result does not establish superiority across all competitive metrics.

The table separates official leaderboard, frontier-harness, and controlled rows because they support different comparisons. AiScientist obtains 81.82 ±0.00 Any Medal% with both backbones and valid submissions on every task, but individual medal tiers are not uniformly best.

MLE-Bench Lite validity, above-median, and medal percentages, reported as mean ± SEM over three runs/seeds.
MLE-Bench Lite validity, above-median, and medal percentages, reported as mean ± SEM over three runs/seeds.

4.4. Long-Horizon Improvement Dynamics

Figure 3 separates the magnitude, breadth, and timing of improvement under GLM-5. AiScientist begins more slowly than AIDE and Codex but continues improving after their mean best-so-far curves largely plateau; its pairwise advantage broadens before the mean score fully overtakes the references.

  • Only 27% of runs reach their final best score by 6 hours, and 44% do so by 12 hours, indicating that much refinement occurs later in the 24-hour budget.
  • The mean normalized validation curve uses ±1 SEM across valid task–seed trajectories. Pairwise lead rate assigns 1 for a win, 0.5 for a tie, and 0 for a loss, with a binomial standard-error approximation.
  • The convergence curve measures time to the final best score observed within the available budget, not convergence to a globally optimal solution.
  • The slower initial progress is consistent with early investment in understanding, planning, and scaffold construction, but the curves do not independently measure the contribution of each activity.

The panels distinguish average improvement from the breadth of paired wins and the time needed to attain the final observed best. Only 27% of runs attain that best by 6 hours and 44% by 12 hours, so early performance does not capture the eventual validation advantage.

Mean best-so-far normalized validation score, pairwise lead rate, and time to final best on MLE-Bench Lite under GLM-5.
Mean best-so-far normalized validation score, pairwise lead rate, and time to final best on MLE-Bench Lite under GLM-5.

4.5. Mechanism Analysis

Figure 4 reports GLM-5 ablations: removing File-as-Bus lowers PaperBench average score by 6.41 points and MLE-Bench Lite Any Medal% by 31.82 percentage points. Valid Submission and Bronze outcomes remain largely intact, while Above Median, Silver, Gold, and Any Medal outcomes decline more strongly.

  • This pattern supports a role for durable evidence in diagnosis and refinement beyond producing a valid submission; the final validity metric alone does not establish when a runnable first pass was obtained.
  • Without File-as-Bus, the hierarchical variant still exceeds the simpler agent by 4.74 PaperBench points, 22.73 Above Median percentage points, and 9.09 Any Medal percentage points.
  • The ablations support contributions from both state continuity and role organization, but do not separately quantify the effects of each schema feature or specialist.

Removing File-as-Bus preserves submission validity but reduces stronger competitive outcomes. The 6.41-point PaperBench loss and 31.82-percentage-point Any Medal loss support durable state as a contributor to refinement; the remaining advantage over the simpler agent also supports role organization, without isolating every component.

GLM-5 mechanism analysis comparing a simpler agent, AiScientist, and AiScientist without File-as-Bus on both benchmarks.
GLM-5 mechanism analysis comparing a simpler agent, AiScientist, and AiScientist without File-as-Bus on both benchmarks.

4.6. Behavioral Analysis on PaperBench

Figure 5 groups PaperBench activity by workflow stage to compare systems with different role structures under GLM-5. AiScientist averages 1112.2 steps versus 834.2 for IterativeAgent, with a particularly large shift toward implementation: 573.5 steps versus 105.5, while auxiliary operations fall from 291.4 to 31.2.

  • Codex averages only 148.4 steps while remaining competitive, demonstrating that longer trajectories are not intrinsically better.
  • AiScientist spends 273.6 steps on experimentation and 129.6 on validation; IterativeAgent spends 217.9 and 187.6, respectively. Contrary to the paper’s broad wording, validation steps do not increase in this comparison.

AiScientist’s 573.5 implementation steps versus IterativeAgent’s 105.5 contrast with its smaller auxiliary allocation, 31.2 versus 291.4. Its validation count is also lower, 129.6 versus 187.6, so the shift is not an increase in every engineering stage. Codex remains competitive with 148.4 total steps, making allocation more informative than trajectory length alone.

Average PaperBench workflow-step allocation by comprehension, implementation, experimentation, validation, auxiliary operations, and total activity under GLM-5.
Average PaperBench workflow-step allocation by comprehension, implementation, experimentation, validation, auxiliary operations, and total activity under GLM-5.

Figure 6 divides tasks into higher- and lower-scoring halves separately for AiScientist and its File-as-Bus ablation under GLM-5. With File-as-Bus, higher-scoring trajectories use 97.2 more validation steps and 72.8 fewer experiment steps; without it, higher-scoring trajectories use 117.8 more implementation steps and 36.8 more experiment steps.

  • The pattern is consistent with successful artifact-mediated trajectories emphasizing validation rather than repeatedly exploring implementation variants.
  • Implementation steps also increase by 23.9 in the higher-scoring File-as-Bus group, so the evidence supports a change in emphasis rather than the absence of additional implementation.
  • These are retrospective associations across different tasks, not controlled interventions on validation effort. Task difficulty and other trajectory differences can affect the observed allocations.

With File-as-Bus, higher-scoring tasks use 97.2 more validation steps and 72.8 fewer experiment steps, while implementation increases modestly by 23.9. Without it, implementation and experimentation increase by 117.8 and 36.8 steps. These task-group associations are consistent with different work patterns but do not establish a causal effect of validation effort or evidence reuse.

Workflow-step differences between higher- and lower-scoring PaperBench task halves for AiScientist and its File-as-Bus ablation under GLM-5.
Workflow-step differences between higher- and lower-scoring PaperBench task halves for AiScientist and its File-as-Bus ablation under GLM-5.

The paper positions AiScientist alongside automated scientific discovery, objective-driven ML engineering, and paper-to-code reproduction systems. Its emphasis is maintaining coherent engineering progress across these stages through hierarchical orchestration and artifact-mediated continuity, rather than introducing a model-training algorithm.

6. Conclusion

The conclusion interprets the benchmark gains and File-as-Bus ablations as evidence that cumulative, inspectable project state contributes to long-horizon research automation. The demonstrated scope is autonomous engineering and refinement on paper-replication and competition tasks under finite budgets, not unrestricted autonomous scientific discovery.

Appendix

  • A. Additional Related Work, including A.1. Automating AI Research and A.2. Multi-Agent Coordination and Long-Horizon Continuity, expands the comparison with discovery systems, objective-driven engineering agents, and reproduction tools. It connects role decomposition in CAMEL, MetaGPT, and ChatDev to documented problems of brittle handoffs, weak verification, and loss of decision-relevant context.
  • B. File-as-Bus Implementation Details specifies role-scoped artifact ownership and access. Table 3 lists paper analyses, agent/prioritized_task.md, submission/, submission/setup/, submission/reproduce.sh, agent/impl_log.md, and agent/exp_log.md; Tier-2 subagents default to read-only access, and the workspace map is refreshed after specialist invocations. Analyses and plans are designated versioned artifacts with conceptually preserved revisions, while runnable files are mutable and their decision histories are recorded in append-only logs.
  • C. Evaluation and Metric Details states that anonymized implementation, evaluation scripts, and experiment artifacts are included in the submitted supplementary material. It defines the anytime mean and SEM over valid trajectories, tie-aware pairwise lead rates, and their binomial standard-error approximation, while specifying the 20-task PaperBench and 22-task MLE-Bench Lite settings.
  • D. Baseline and Reference Details separates controlled comparisons from contextual calibration. BasicAgent and IterativeAgent share the PaperBench task and grading setup, whereas Codex/GPT-5.5 xhigh differs in both harness and backbone; official MLE-Bench Lite leaderboard entries may also differ in model and reporting protocol.
  • E. Supplementary Analyses includes E.1. Delegation Patterns on PaperBench and E.2. Budget and Medal Outcomes on MLE-Bench Lite. Figure 7 shows delegation across the task suite, with mean Orchestrator and subagent tool-call counts of 94.7 and 1037.0. Figure 8 shows that 29% of medal runs have converged by 6 hours and 49% by 12 hours; it describes convergence within medal runs, not the fraction of all tasks earning medals at each time.

Brief Thoughts

The clearest contribution is a collaboration contract for persistent evidence, supported by an ablation in which validity remains intact while competitive outcomes deteriorate. Table 3 assigns responsibilities and artifact lifecycles; Figures 7 and 8 show that delegation is routinely used and that many medal runs reach their final best only in the later part of the budget.

Writer and reader assignments make artifacts explicit interfaces between stages. Analyses and plans are designated versioned artifacts, runnable files remain mutable, and diagnostic logs are append-only. This contract separates current executable state from its recorded rationale, although the accompanying text describes revision preservation conceptually rather than specifying a versioning mechanism.

File-as-Bus artifact schema specifying purposes, update modes, writers, and readers for analysis, planning, execution, and diagnostic records.
File-as-Bus artifact schema specifying purposes, update modes, writers, and readers for analysis, planning, execution, and diagnostic records.

Delegated tool calls occur throughout the PaperBench suite rather than in only a few tasks. Mean subagent activity is 1037.0 calls per task versus 94.7 Orchestrator calls, documenting substantial delegated execution without by itself showing that delegation causes higher scores.

Orchestrator and subagent tool calls across PaperBench tasks under GLM-5.
Orchestrator and subagent tool calls across PaperBench tasks under GLM-5.

Among medal runs, 29% have converged by 6 hours and 49% by 12 hours. The curve supports late-budget refinement within this subset; it does not measure the percentage of all tasks earning medals at each time or directly estimate outcomes under separately imposed shorter budgets.

Cumulative fraction of converged medal runs over the 24-hour MLE-Bench Lite budget under GLM-5.
Cumulative fraction of converged medal runs over the 24-hour MLE-Bench Lite budget under GLM-5.

The evidence supports the design within the tested engineering setting, but several limits remain. PaperBench’s approximately $832 grading cost constrains repeated evaluation, its main results lack the three-run uncertainty reporting used for MLE-Bench Lite, and workflow-step associations cannot establish causality. The Codex comparison also changes both model and harness, so its score difference is not an architectural ablation.