Can AI agents conduct open-ended AI research? Early evidence from two case studies | Summary
29 Jul 2026 | Paper Review AI Research Agents Open-World Evaluation LLM EvaluationContents
- Summary
- 2 Shadow evaluations: A new method for measuring progress towards automating AI research
- 3 The task and the research setup
- 4 Results
- 4.1 The agents did not produce research at the caliber of a top ML conference
- 4.3 The agents committed to unpromising approaches too quickly
- 4.4 The agents could not make good use of feedback from AI reviews
- 4.5 The agents were capable of all of the engineering steps required to conduct the research
- 5 Log analysis reveals five failure modes
- 6 A robustness experiment with Codex and GPT-5.6 Sol Ultra reproduced these failure modes
- 7 Limitations
- Appendix
- Brief Thoughts
This article explains the key points of Can AI agents conduct open-ended AI research? Early evidence from two case studies.
- 2026-07-29 (arXiv)
- Kirgis, Peter, Kapoor, Sayash, Schwartz, Andrew, Rabanser, Stephan, Africa, David, Voudouris, Konstantinos, Nguyen, Viet, Pilditch, Toby, Dubois, Magda, Coppock, Harry, et al.
- Princeton University, Cornflower Labs, UK AI Security Institute, University of Toronto, Independent, UC Berkeley, Georgetown University (CSET), Johns Hopkins University, Golden Gate Institute for AI, AI Digest, Stanford University
- Paper
Summary
- The paper introduces shadow evaluations for testing whether frontier agents can conduct open-ended AI research. An agent receives the central question of a high-quality unpublished paper, works without access to its findings, and is assessed by the original authors as a conference reviewer; in two six-day case studies, agents completed the engineering work but did not produce research judged suitable for a top ML conference.
Original-paper authors scored Personas at 2/6 overall and TabPFN at 1/6, with low quality, clarity, and significance ratings in both cases. The table grounds the central outcome in expert assessment rather than an automated metric.
2 Shadow evaluations: A new method for measuring progress towards automating AI research
Shadow evaluations are proposed as a complement to verifier-scored benchmarks and blind peer review. They use authentic, uncontaminated research questions and evaluators with unusually deep knowledge of the target problem, but trade scalability and objective ground truth for non-blind expert judgment and a small sample size.
The comparison places shadow evaluations between automatically verified tasks with narrow objectives and end-to-end research evaluated by human review. It supports the authors’ claim that shadow evaluation complements rather than replaces existing evaluation paradigms.
3 The task and the research setup
The study used two unpublished NeurIPS 2026 submissions: a Personas question about decomposing and controlling LLM personas through weight-space interventions, and a TabPFN question about deployment-time detection of harmful distribution shift. Claude Opus 4.8 with extra-high reasoning ran in OpenClaw with a Linux VM, web access, subagents, GPU resources, AI reviewing tools, six days of wall-clock time, $3,000 in API credits, and a later 24-hour extension.
4 Results
The original authors rejected both agent-generated papers. Personas received an overall score of 2/6 and TabPFN received 1/6; reviews cited poorly motivated data and experimental choices, unsupported conclusions, limited significance, and unclear presentation, while recognizing some reasonable initial hypotheses and minor potentially relevant findings.
4.1 The agents did not produce research at the caliber of a top ML conference
Neither submission met the reviewers’ standard for a top ML conference. The TabPFN reviewer characterized a broad conclusion from a few unsuccessful white-box signals as a “proof by example” fallacy, while the Personas reviewer questioned post-hoc choices of traits, datasets, and lexical measures, despite considering the weight-difference approach relatively original.
4.3 The agents committed to unpromising approaches too quickly
Both agents developed plausible initial directions but rejected ambitious hypotheses after small-scale or underpowered tests, then converged on weaker negative-result narratives. The Personas agent planned 42 hours of open-ended exploration but settled on a method after five hours, while the TabPFN agent committed to its headline direction roughly 40 hours before its planned exploration deadline.
The trajectories show that the Personas and TabPFN agents used $1,130/$3,000 and $1,235/$3,000 of API budget, respectively, even after deadline extensions. The annotations link early hypothesis retirement, weak-reject reviews, debugging successes, and premature completion to the paper’s diagnoses of poor resource awareness and backtracking.
Planned-versus-realized milestones document premature convergence in both studies: Personas exploration ended after five hours despite a 42-hour plan, and TabPFN committed to its headline finding 40 hours early. The figure provides direct evidence that agents moved to drafting before adequately exploring alternatives.
4.4 The agents could not make good use of feedback from AI reviews
Self-reviews consistently returned reject or weak-reject verdicts, and external tools surfaced several concerns later raised by human experts. Rather than redesigning experiments or reconsidering their premises, the agents generally narrowed claims and added caveats, while placing too much weight on comparatively lenient acceptance recommendations; the authors frame this as suggestive, but not conclusive, evidence of a generator–verifier gap.
The review histories show that self-reviews repeatedly issued reject or weak-reject verdicts, while the Stanford Agentic Reviewer accepted early drafts and human experts ultimately rejected both papers. This pattern supports the claim that agents identified critiques but did not appropriately prioritize or act on the most consequential ones.
Human experts and agents’ final self-reviews overlap on specific weaknesses, including unsupported generalization, unclear writing, hand-picked traits, weak measures, and inadequate validation. The authors use this agreement to argue that the problem was not only detecting flaws, but deciding which flaws demanded a fundamental change of approach.
4.5 The agents were capable of all of the engineering steps required to conduct the research
The agents independently performed substantial engineering work, including literature review, GPU-environment debugging, hundreds of experiments and robustness checks, web- and email-based retrieval of reviews, and compilation of complete LaTeX manuscripts. They resolved nearly all environmental barriers, although an OpenClaw bug required a manual patch; both finished papers nevertheless exceeded the NeurIPS page limit and had weak presentation.
The side-by-side pages show the effect of a human-requested readability pass. Although readability improved, the intervention and surrounding results illustrate that research engineering competence did not produce conference-ready communication or compliance with format requirements.
5 Log analysis reveals five failure modes
Log analysis identifies five recurring failure modes: insufficient judgment about the standard for publishable research, inability to creatively repair weak designs, ineffective project-level backtracking, poor awareness of time and budget, and instruction drift. The agents used less than half of their API budgets despite monitoring them, failed to maintain requirements such as exploration time and paper-length limits, and made local experimental revisions without restarting unpromising projects.
6 A robustness experiment with Codex and GPT-5.6 Sol Ultra reproduced these failure modes
A robustness run repeated the TabPFN task with GPT-5.6 Sol at ultra reasoning in Codex, using the same time and API budgets and a loop that re-engaged the agent after review. It reproduced central shortcomings—underpowered experiments, no novel contribution, and malformed output—while using a real-world shifted dataset, registering hypotheses, and reporting a negative finding; it exhausted the $3,000 API budget in just over two days with nearly 100 hours remaining.
7 Limitations
The authors emphasize that five runs and two questions cannot establish broad capability limits, and that non-blind original-author review, question selection, scaffold choices, and interpretive discretion may bias the results. OpenClaw session resets lost accumulated context, the study did not test Anthropic’s strongest model, and six days were far shorter than the original authors’ projects; however, the authors argue that unused resources and reviewers’ objections to research judgment make additional time unlikely to change the main outcome.
Appendix
- The appendix specifies the two target questions, supplies full reviews from the original authors, reports a pre-experiment survey of twelve collaborators, surveys prior autonomous AI R&D evaluations, and diagrams the CRUX 2 OpenClaw scaffold. The reviews directly document the reported judgments: the TabPFN review assigns 1/4 quality and a 1/6 strong reject, while the Personas review assigns 2/4 quality and a 2/6 reject.
Brief Thoughts
The paper’s central contribution is an evaluation design in which agents address unpublished research questions and are assessed by domain experts familiar with the underlying problems. Its conclusions remain provisional: the experiments give detailed evidence of failure on two empirical ML projects and one robustness replication, but the small, non-blind sample cannot establish how broadly the observed failure modes generalize across models, scaffolds, or research styles.