Gorio Tech Blog search

Can AI agents conduct open-ended AI research? Early evidence from two case studies | Summary

|

Contents

This article explains the key points of Can AI agents conduct open-ended AI research? Early evidence from two case studies.

  • 2026-07-29 (arXiv)
  • Kirgis, Peter, Kapoor, Sayash, Schwartz, Andrew, Rabanser, Stephan, Africa, David, Voudouris, Konstantinos, Nguyen, Viet, Pilditch, Toby, Dubois, Magda, Coppock, Harry, et al.
  • Princeton University, Cornflower Labs, UK AI Security Institute, University of Toronto, Independent, UC Berkeley, Georgetown University (CSET), Johns Hopkins University, Golden Gate Institute for AI, AI Digest, Stanford University
  • Paper

Read this article in Korean


Summary

  • The paper introduces shadow evaluations for testing whether frontier agents can conduct open-ended AI research. An agent receives the central question of a high-quality unpublished paper, works without access to its findings, and is assessed by the original authors as a conference reviewer; in two six-day case studies, agents completed the engineering work but did not produce research judged suitable for a top ML conference.

Original-paper authors scored Personas at 2/6 overall and TabPFN at 1/6, with low quality, clarity, and significance ratings in both cases. The table grounds the central outcome in expert assessment rather than an automated metric.

Summary of the original paper authors’ reviews of the agents’ submitted work, including criterion scores, overall verdicts, confidence, and expert comments.
Summary of the original paper authors’ reviews of the agents’ submitted work, including criterion scores, overall verdicts, confidence, and expert comments.

2 Shadow evaluations: A new method for measuring progress towards automating AI research

Shadow evaluations are proposed as a complement to verifier-scored benchmarks and blind peer review. They use authentic, uncontaminated research questions and evaluators with unusually deep knowledge of the target problem, but trade scalability and objective ground truth for non-blind expert judgment and a small sample size.

The comparison places shadow evaluations between automatically verified tasks with narrow objectives and end-to-end research evaluated by human review. It supports the authors’ claim that shadow evaluation complements rather than replaces existing evaluation paradigms.

Selected evaluations and demonstrations of experiments studying AI research and development, organized by agent task and evaluator.
Selected evaluations and demonstrations of experiments studying AI research and development, organized by agent task and evaluator.

3 The task and the research setup

The study used two unpublished NeurIPS 2026 submissions: a Personas question about decomposing and controlling LLM personas through weight-space interventions, and a TabPFN question about deployment-time detection of harmful distribution shift. Claude Opus 4.8 with extra-high reasoning ran in OpenClaw with a Linux VM, web access, subagents, GPU resources, AI reviewing tools, six days of wall-clock time, $3,000 in API credits, and a later 24-hour extension.

4 Results

The original authors rejected both agent-generated papers. Personas received an overall score of 2/6 and TabPFN received 1/6; reviews cited poorly motivated data and experimental choices, unsupported conclusions, limited significance, and unclear presentation, while recognizing some reasonable initial hypotheses and minor potentially relevant findings.

4.1 The agents did not produce research at the caliber of a top ML conference

Neither submission met the reviewers’ standard for a top ML conference. The TabPFN reviewer characterized a broad conclusion from a few unsuccessful white-box signals as a “proof by example” fallacy, while the Personas reviewer questioned post-hoc choices of traits, datasets, and lexical measures, despite considering the weight-difference approach relatively original.

4.3 The agents committed to unpromising approaches too quickly

Both agents developed plausible initial directions but rejected ambitious hypotheses after small-scale or underpowered tests, then converged on weaker negative-result narratives. The Personas agent planned 42 hours of open-ended exploration but settled on a method after five hours, while the TabPFN agent committed to its headline direction roughly 40 hours before its planned exploration deadline.

The trajectories show that the Personas and TabPFN agents used $1,130/$3,000 and $1,235/$3,000 of API budget, respectively, even after deadline extensions. The annotations link early hypothesis retirement, weak-reject reviews, debugging successes, and premature completion to the paper’s diagnoses of poor resource awareness and backtracking.

Annotated summaries of the agents’ time, API, and GPU compute expenditures.
Annotated summaries of the agents’ time, API, and GPU compute expenditures.

Planned-versus-realized milestones document premature convergence in both studies: Personas exploration ended after five hours despite a 42-hour plan, and TabPFN committed to its headline finding 40 hours early. The figure provides direct evidence that agents moved to drafting before adequately exploring alternatives.

Comparison between the agents’ planned and realized wall-clock budgets for milestones in the research process.
Comparison between the agents’ planned and realized wall-clock budgets for milestones in the research process.

4.4 The agents could not make good use of feedback from AI reviews

Self-reviews consistently returned reject or weak-reject verdicts, and external tools surfaced several concerns later raised by human experts. Rather than redesigning experiments or reconsidering their premises, the agents generally narrowed claims and added caveats, while placing too much weight on comparatively lenient acceptance recommendations; the authors frame this as suggestive, but not conclusive, evidence of a generator–verifier gap.

The review histories show that self-reviews repeatedly issued reject or weak-reject verdicts, while the Stanford Agentic Reviewer accepted early drafts and human experts ultimately rejected both papers. This pattern supports the claim that agents identified critiques but did not appropriately prioritize or act on the most consequential ones.

Summary of the paper reviews from subagents, external AI reviews, and the expert human reviews.
Summary of the paper reviews from subagents, external AI reviews, and the expert human reviews.

Human experts and agents’ final self-reviews overlap on specific weaknesses, including unsupported generalization, unclear writing, hand-picked traits, weak measures, and inadequate validation. The authors use this agreement to argue that the problem was not only detecting flaws, but deciding which flaws demanded a fundamental change of approach.

Detailed comparison of the human experts’ and agents’ final reviews of the agents’ work.
Detailed comparison of the human experts’ and agents’ final reviews of the agents’ work.

4.5 The agents were capable of all of the engineering steps required to conduct the research

The agents independently performed substantial engineering work, including literature review, GPU-environment debugging, hundreds of experiments and robustness checks, web- and email-based retrieval of reviews, and compilation of complete LaTeX manuscripts. They resolved nearly all environmental barriers, although an OpenClaw bug required a manual patch; both finished papers nevertheless exceeded the NeurIPS page limit and had weak presentation.

The side-by-side pages show the effect of a human-requested readability pass. Although readability improved, the intervention and surrounding results illustrate that research engineering competence did not produce conference-ready communication or compliance with format requirements.

Comparison of the submitted papers before and after a human intervention which instructed a final pass from the agent for readability.
Comparison of the submitted papers before and after a human intervention which instructed a final pass from the agent for readability.

5 Log analysis reveals five failure modes

Log analysis identifies five recurring failure modes: insufficient judgment about the standard for publishable research, inability to creatively repair weak designs, ineffective project-level backtracking, poor awareness of time and budget, and instruction drift. The agents used less than half of their API budgets despite monitoring them, failed to maintain requirements such as exploration time and paper-length limits, and made local experimental revisions without restarting unpromising projects.

6 A robustness experiment with Codex and GPT-5.6 Sol Ultra reproduced these failure modes

A robustness run repeated the TabPFN task with GPT-5.6 Sol at ultra reasoning in Codex, using the same time and API budgets and a loop that re-engaged the agent after review. It reproduced central shortcomings—underpowered experiments, no novel contribution, and malformed output—while using a real-world shifted dataset, registering hypotheses, and reporting a negative finding; it exhausted the $3,000 API budget in just over two days with nearly 100 hours remaining.

7 Limitations

The authors emphasize that five runs and two questions cannot establish broad capability limits, and that non-blind original-author review, question selection, scaffold choices, and interpretive discretion may bias the results. OpenClaw session resets lost accumulated context, the study did not test Anthropic’s strongest model, and six days were far shorter than the original authors’ projects; however, the authors argue that unused resources and reviewers’ objections to research judgment make additional time unlikely to change the main outcome.

Appendix

  • The appendix specifies the two target questions, supplies full reviews from the original authors, reports a pre-experiment survey of twelve collaborators, surveys prior autonomous AI R&D evaluations, and diagrams the CRUX 2 OpenClaw scaffold. The reviews directly document the reported judgments: the TabPFN review assigns 1/4 quality and a 1/6 strong reject, while the Personas review assigns 2/4 quality and a 2/6 reject.

Brief Thoughts

The paper’s central contribution is an evaluation design in which agents address unpublished research questions and are assessed by domain experts familiar with the underlying problems. Its conclusions remain provisional: the experiments give detailed evidence of failure on two empirical ML projects and one robustness replication, but the small, non-blind sample cannot establish how broadly the observed failure modes generalize across models, scaffolds, or research styles.