The Last Human-Written Paper: Agent-Native Research Artifacts | Summary
27 Apr 2026 | Paper Review AI Research Agents Research Reproducibility Agentic Skills Agent's MemoryContents
- Summary
- 1. Introduction
- 2. The ARA Protocol
- 3. The Live Research Manager
- 4. The ARA Compiler
- 5. ARA Verification and Review
- 6. The (Human+AI)² Research Network
- 7. Evaluation
- 8. Related Work
- 9. Future Work
- 10. Limitations
- 11. Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of The Last Human-Written Paper: Agent-Native Research Artifacts.
- 2026-04-27 (arXiv)
- Liu, Jiachen, Pei, Jiaxin, Huang, Jintao, Si, Chenglei, Qu, Ao, Tang, Xiangru, Lu, Runyu, Chen, Lichang, Bai, Xiaoyan, Zheng, Haizhong, et al.
- University of Michigan, Stanford University, Ohio State University, MIT, Yale University, Meta Superintelligence Labs, University of Chicago, Carnegie Mellon University, University of Washington, University of Toronto, NVIDIA, Meta, New York University, Nanyang Technological University, Orchestra Research, Harvard University, LinkedIn, UIUC, Arizona State University, Stony Brook University, University of Hong Kong, Boston College, Portland State University, National University of Singapore, Cornell University
- Paper
- Github
- Project Page
Summary
- The Last Human-Written Paper: Agent-Native Research Artifacts proposes the Agent-Native Research Artifact (ARA), a file-system protocol for making research directly operable by AI agents. It separates scientific logic, executable specifications, exploration history, and raw evidence, with explicit links between claims, experiments, code, and outputs.
- Three mechanisms populate and verify the artifact: a Live Research Manager captures researcher–agent sessions, an ARA Compiler imports legacy research sources, and the ARA Seal checks structural integrity, argumentative rigor, and directional reproducibility. Human review remains responsible for significance, novelty, problem formulation, and ethical judgment.
- Across 450 knowledge-extraction questions, ARA achieves 93.7% accuracy versus 72.4% for the conventional baseline; across 150 reproduction subtasks, the paper reports difficulty-weighted success of 64.4% versus 57.4%. ARA additionally incorporates expert reproduction rubrics or prior exploration transcripts, so these comparisons do not isolate representation alone, and the reproduction appendix contains inconsistencies between its summaries and tabulated results.
- On five RE-Bench extension tasks, preserved failure knowledge helps agents make earlier useful moves, but ARA finishes ahead on only three tasks under Claude Sonnet 4.6. The remaining trajectories suggest that prior heuristics can anchor or fragment search; deployment security, privacy controls, and generalization beyond computational ML remain unresolved.
1. Introduction
The introduction distinguishes the Storytelling Tax—the loss of failed experiments, rejected hypotheses, and branching decisions—from the Engineering Tax—the gap between reviewer-sufficient prose and agent-sufficient execution specifications. Figures 1–2 motivate preserving the research object rather than only its successful narrative, while Figure 3 quantifies missing reproduction information.
- Across 8,921 PaperBench reproduction requirements from 23 papers, only 45.4% are judged fully specified in the PDF; code development has 37.3% sufficient coverage, and missing hyperparameters account for 26.2% of reported gap labels.
- The exploration-cost analysis assigns 59.2% of tokens and 90.2% of dollar cost to below-reference runs. Appendix E.3 clarifies that the 24,008-run corpus spans 21 models and 228 tasks, not exclusively RE-Bench, and that unsuccessful exploration becomes an avoidable downstream cost only when successors must repeat it.
- The proposal makes ARA canonical and treats papers as compiled views. The problem is attributed to documentation conventions, not simply to the PDF file format.
The comparison identifies decision history, failed experiments, configuration rationale, and code bindings as information that a conventional narrative can omit. ARA’s four linked layers are presented as the alternative operational object; the diagram states the design motivation rather than measuring preservation fidelity.
The branching tree includes wrong losses, out-of-memory failures, divergence, and reward hacking, whereas the published sequence retains only a successful path. This illustrates why a final-method description cannot tell successors which alternatives were already explored.
Only 45.4% of the 8,921 requirements are judged fully specified, with code development showing the lowest sufficient coverage at 37.3%. Missing hyperparameters, vague descriptions, and cross-reference-only specifications dominate the reported gap-label distribution, motivating explicit execution specifications.
2. The ARA Protocol
The ARA protocol defines a file-system ontology for research knowledge rather than a new model or training algorithm. Its four layers answer distinct operational questions: what the contribution is, how to reproduce it, what prior exploration established, and which empirical outputs support its claims.
2.1. Design Philosophy
The organizing principle is Knowledge over Narrative: evolving research knowledge is the primary object, and prose is a derived presentation. Figure 4 specifies independently queryable directories, while Figure 5 illustrates cross-layer bindings that let an agent follow a claim to its experiment, implementation, evidence, and relevant failures.
- Progressive disclosure limits context use by loading task-relevant files instead of repeatedly scanning an entire paper or repository.
- The protocol favors fact-dense statements and provenance over narrative framing. High-fidelity preservation is a design requirement, not a demonstrated guarantee that every compiled artifact is information-lossless.
The directory tree assigns distinct retrieval locations to problem definitions, claims, experiment plans, algorithms, configurations, implementations, exploration history, and outputs. It makes progressive disclosure concrete: an agent can use PAPER.md to select relevant files without loading the entire artifact.
Arrows connect a scientific claim to an experiment, implementation, heuristic, and raw result, while the exploration graph retains a rejected branch. The mechanism is traceable support across layers, not simply placing documents in separate folders.
2.2. ARA Architecture
PAPER.md supplies YAML frontmatter and a layer index intended to support relevance triage in approximately 500 tokens. The architecture separates experiment plans from exact outcomes so that the proposed verification workflow can withhold evidence while an agent attempts reproduction.
- The Cognitive Layer (/logic) contains problem.md, solution/, claims.md, experiments.md, and related_work.md. Claims carry falsification criteria and proof pointers; typed literature dependencies express imports, bounds, and baselines.
- The Physical Layer (/src) uses kernel mode for separable algorithmic contributions and repository mode when engineering is itself the contribution. configs/ records parameters and rationale, while environment.md records dependencies, hardware, and seeds.
- The Exploration Graph (/trace) stores question, decision, experiment, dead_end, and pivot nodes in exploration_tree.yaml; also_depends_on represents converging dependencies. The Evidence Layer (/evidence) stores exact results, logs, and diagnostics.
- Separating experiment logic, outcomes, and recorded accept/reject decisions could also support agent training environments. Artifact sufficiency is defined relative to a sufficiently capable coding agent reproducing the core claim without human intervention or external context, not as a guarantee for present-day agents.
3. The Live Research Manager
The Live Research Manager addresses the burden of manually authoring structured research records. It treats researcher–agent conversations, tool outputs, experiment results, and code changes as source material for a living ARA; Figure 6 shows capture at session boundaries rather than interruption during active work.
3.1. Design Principles
The manager follows three principles: silent and framework-independent integration, faithful epistemic provenance, and comprehensive trajectory capture. Its natural-language agent skill uses existing file and shell tools, tags the origin of research events, and preserves revisions through version control.
- Provenance distinguishes user, ai-suggested, ai-executed, and user-revised events. An ai-suggested event is not automatically promoted without researcher confirmation.
- Unsettled observations enter /staging/ rather than being forced immediately into claims or decisions. The claimed negligible documentation burden assumes that relevant research activity already occurs within agent-mediated sessions; capture overhead is not separately benchmarked.
3.2. System Design
At session close, the Context Harvester extracts research-significant activity, the Event Router assigns types and provenance, and the Maturity Tracker promotes settled observations. Table 1 lists seven event types and their structured payloads; trace events accumulate at session boundaries, while claims and specifications crystallize at milestones.
- Closure signals include topic abandonment, explicit affirmation, empirical resolution, and commitment in a downstream artifact. Contradictions are retained and flagged instead of silently overwriting prior entries.
- The artifact carries cross-session memory through session records, a session index, staged observations, and trace/pm_reasoning_log.yaml. Although the overview mentions a structured briefing, the detailed continuity mechanism surfaces relevant context when needed rather than requiring an unsolicited opening briefing.
- The implementation is available at https://github.com/AmberLJC/Agent-Native-Research-Artifact.
Capture occurs after an active researcher–agent session through harvesting, routing, and maturity tracking. The feedback path from the persistent artifact to a later session illustrates how the manager can remain stateless while the research record supplies continuity.
The seven event types distinguish trajectory entries from scientific claims, implementation heuristics, and unclassified observations. Requiring a dead end to include its hypothesis, failure mode, and lesson turns an unsuccessful run into potentially reusable knowledge rather than merely an archived log.
4. The ARA Compiler
The ARA Compiler provides backward compatibility for research that was not captured during development. It accepts PDFs, repositories, datasets, expert rubrics, and experimental trajectories, then reconstructs a canonical artifact; Figure 7 separates in-loop Level 1 validation from downstream Levels 2–3 verification.
4.1. Design Principles
The compiler’s principles are universal input with canonical output, high-fidelity preservation, and knowledge lineage rather than flat extraction. A PDF-only input may produce a structurally valid artifact with code stubs, whereas repositories, rubrics, and transcripts can supply implementation and process knowledge absent from the paper.
- Preservation requires retaining source numerical results, configurations, architectural details, and negative findings rather than summarizing them away.
- The central reconstruction task is to recover claim–experiment–evidence–code connections implicit across prose, figures, appendices, and repository comments. Structural conformance does not establish that these reconstructed connections are scientifically correct.
4.2. Compiler Implementation
Compilation proceeds through Semantic Deconstruction, Cognitive Mapping, Physical Grounding, and Exploration Graph Extraction. Level 1 diagnostics drive a generate→validate→fix loop, described as typically requiring 2–3 rounds; auxiliary sources are routed to the layers they support.
- Physical Grounding replaces stubs with available implementations and reconciles code against paper claims, recording tacit implementation knowledge as provenance-tagged heuristics.
- When prior ARAs are available, collective inference can propose common domain heuristics absent from the current paper. These are tagged collective_inference so inferred candidates remain distinguishable from source-stated facts; the reported evaluations do not isolate this mechanism.
Multiple source types feed four top-down compilation stages, with Level 1 validation providing repair feedback. Source enrichment is part of the method, while structural conformance during compilation is separated from later argumentative and empirical verification.
5. ARA Verification and Review
ARA-native verification separates deterministic structural checks, rubric-based assessment of scientific support, and expert judgment. The ARA Seal is the proposed credential for automated checks; the review pipeline consumes its reports rather than replacing human evaluation of research value.
5.1. Design Principles
The review system seeks to automate mechanical verification and make reproducibility a foundational submission requirement. Structural failure blocks advancement, argumentative critique precedes costly execution, and human reviewers receive the audit trail to evaluate significance, novelty, and contested findings.
5.2. The ARA Seal: Machine-Verifiable Research Credentials
The text defines three automated ARA Seal levels: Structural Integrity, Argumentative Rigor, and Execution Reproducibility. Figure 8 also includes a Human Judgment card labeled Level 4, although the surrounding protocol reserves human judgment for review Stage 3 rather than a fourth automated Seal level.
- Level 1 checks directories, required fields, schemas, minimum counts, and reference resolution deterministically.
- Level 2 uses a Rigor Auditor without code execution or external consultation. It scores evidence relevance, falsifiability quality, scope calibration, argument coherence, exploration integrity, and methodological rigor on anchored 1–5 rubrics, emitting severity-ranked findings with evidence spans. Rubric anchoring does not make its judgments deterministic.
- Level 3 prioritizes central claims and performs budget-limited, scaled-down directional checks rather than exact full-scale replication. The proposed protocol withholds /evidence/ and marks unaffordable claims unverified; optional full-scale reproductions would update the certificate later.
- A Seal Certificate is specified as a signed record of artifact ID, achieved level, timestamp, environment hash, and per-claim outcomes. Production sandboxing and granular access controls remain unimplemented.
The cards distinguish inexpensive structural checks, rubric-based argumentative assessment, costly execution checks, and human judgment. The graphic labels Human Judgment as Level 4, whereas the text defines three automated Seal levels; its sandboxed-execution depiction is also a proposed property not implemented in the current system.
5.3. Three-Stage Review Pipeline
Figure 9 organizes review into Conceptual Verification (minutes), Empirical Verification (hours–days), and Human Review (days–weeks). Stage 1 combines structural checks with a rigor report; Stage 2 attempts prioritized directional reproduction under a venue compute budget; Stage 3 evaluates contribution quality and adjudicates disputes.
- Advisory diagnostics include missing dead-end documentation, incomplete baselines, and gaps in experiment–claim coverage. These are distinguished from Seal checks that gate advancement.
- Empirical review would also examine ablations, claimed generality, and undocumented code heuristics. Human findings are linked to specific ARA components, but the paper does not experimentally measure savings in reviewer time.
The funnel places conceptual and empirical verification before expert review and attaches machine-generated reports to the submission. It depicts an intended division of labor, but the paper does not test whether this pipeline reduces reviewer workload or improves human acceptance decisions.
6. The (Human+AI)² Research Network
The (Human+AI)² Research Network composes capture, compilation, verification, and exchange around a canonical artifact. Figure 10 shows researcher-directed agents using /submit, /retrieve, and /fork; downstream teams would retain parent attribution and submit structured changes for review.
- Consumer agents could render the artifact as a paper, slides, video, demonstration, or grounded dialogue according to the reader’s needs.
- Git-like publication, executable contribution diffs, and a shared scientific commons are ecosystem proposals, not outcomes established by the benchmark experiments.
Researchers communicate through agents that submit, retrieve, and fork shared artifacts; agents may also coordinate directly. The diagram extends the artifact protocol into a network proposal rather than presenting an experimentally evaluated multi-team research system.
7. Evaluation
Evaluation tests knowledge extraction, reproduction, and open-ended extension, followed by the automated review machinery. The agent and task are held fixed within comparisons, but the ARA condition combines structured organization with source enrichment that conventional published bundles generally lack.
7.1. Datasets and Motivation
Table 2 distinguishes the complementary benchmarks: PaperBench supplies 23 papers and 8,921 expert reproduction requirements, while RE-Bench supplies seven R&D tasks and recorded agent trajectories. Understanding covers all 30 targets; reproduction uses 15 papers with repositories, and extension uses five tasks with usable failure records.
- The PaperBench baseline receives the PDF and its repository when available; ARA compilation additionally uses the expert rubric. Configuration-recovery gains can therefore reflect both centralized organization and expert-supplied information.
- RE-Bench tasks have no original research papers, so the baseline uses a synthesized polished writeup plus official source. ARA preserves additional MALT exploration knowledge, making this a comparison with a deliberately conventional, success-focused presentation rather than an information-matched alternative.
- The empirical scope is computational ML. The proposed applicability to other digitalized research is not tested, and physical-world interventions or theoretical contributions without executable code remain outside the demonstrated scope.
PaperBench contributes configuration knowledge through expert rubrics, while RE-Bench contributes exploration knowledge through agent transcripts. These additional sources explain why the evaluation measures enriched ARAs rather than a strictly information-matched change in formatting.
7.2. Knowledge Extraction from ARA
Knowledge extraction uses 450 questions: 300 Category A fidelity questions, 115 Category B configuration/detail questions, and 35 Category C failure-knowledge questions. Independent Claude Sonnet 4.6 contexts answer each question, and a Claude Opus 4.6 judge assigns 1.0, 0.5, or 0.0 against a gold reference; question pools generated from both representations are merged and deduplicated.
- Table 3 reports overall accuracy of 93.7% versus 72.4%, with 114.0K versus 109.1K tokens per question. Category A reaches 95.6% versus 80.8%, Category B 92.6% versus 67.8%, and Category C 81.4% versus 15.7%.
- The proposed mechanisms differ: indexed lookup aids fidelity retrieval, centralized configurations aid detail recovery, and preserved transcripts supply failure knowledge unavailable to the baseline. These explanations are consistent with the results but are not separated by ablations.
- The reported 12% token reduction applies to PaperBench Category A (86.3K versus 97.7K), not the complete Category A aggregate (84.6K versus 88.5K). Appendix E reports McNemar χ² = 95.15 and p < 10⁻¹⁰ overall, while its difficulty-stratum counts do not reconcile with the stated total.
The largest accuracy gap occurs on failure knowledge, 81.4% versus 15.7%, where the baseline lacks the relevant trajectory source. Overall token usage is slightly higher for ARA, and retrieval savings are concentrated in fidelity questions rather than uniform across categories.
7.3. Reproduction from ARA
Reproduction comprises 150 subtasks and 1,743 rubric requirements across 15 papers, with 50 easy, 49 medium, and 51 hard tasks. Both conditions use Claude Sonnet 4.6, the same task prompts, and per-paper budgets of 14–20M tokens with cache reads discounted to 10%; a blinded Claude Opus 4.6 judge scores requirements as yes, partial, or no.
- The primary metric weights easy, medium, and hard tasks at 1:2:3. The paper reports difficulty-weighted success of 64.4% for ARA and 57.4% for PDF + GitHub, Wilcoxon p = 0.028, and a paper-level win/tie/loss summary of 8/5/2. These summaries should be treated as reported results because Appendix F does not fully reconcile them with Table 11.
- Figure 11 reports 85.1% versus 80.2% on easy tasks, 68.5% versus 62.9% on medium tasks, and 54.5% versus 46.0% on hard tasks. The widening gap is consistent with operational specifications helping more demanding execution, although substantial hard-task requirements remain unsatisfied.
- The fre case illustrates a possible execution benefit: the ARA agent rebuilt JAX code in PyTorch, used 1.8 GB rather than 30.8 GB GPU memory, and trained 17 models across three domains, whereas the baseline completed three training attempts.
- Expected numbers are masked in prompts, but the listed ARA inputs include evidence/. This is weaker than the Seal’s proposed complete evidence withholding. Fabrication occurred in one ARA run and two baseline runs, so neither condition reliably prevents it.
The reported ARA advantage increases from 4.9 percentage points on easy tasks to 5.6 on medium and 8.5 on hard tasks. This is consistent with greater benefit where execution depends on detailed configurations, although substantial hard-task requirements remain unsatisfied and the appendix does not transparently reconcile the aggregates with its per-paper table.
7.4. Extension from ARA
Extension tests triton_cumsum, rust_codecontests, nanogpt_chat_rl, fix_embedding, and restricted_mlm using identical starting code and a nominal 8 h SLURM wall-clock plus $50 API-spend cap per run. A direction-aware, per-attempt filter excludes MALT scoring attempts that beat the reference; ARA retains below-reference partial successes, failed approaches, and derived heuristics.
- Figure 12 shows that ARA finishes ahead on rust_codecontests, nanogpt_chat_rl, and fix_embedding under Sonnet 4.6, while the paper agent wins on triton_cumsum and restricted_mlm. The case studies link early strategy choices to specific trace entries, but a small set of trajectories cannot establish their average causal effect.
- On rust_codecontests, ARA pursues a hand-coded Rust library within approximately 10 min; the paper agent independently rediscovers that direction at 395 min after prolonged prompt engineering. On fix_embedding, the paper agent retries permutation recovery after already observing its failure.
- The Sonnet 4.6 paper agent overtakes on triton_cumsum through an int8 kernel redesign absent from the trace, and on restricted_mlm through sustained tuning of ConvMLMDilated. ARA instead remains anchored to a prior kernel strategy or spreads effort across named alternatives.
- Paired Sonnet 4.5 comparisons reverse both outcomes: ARA scores 0.27 versus 0.64 on triton_cumsum and 0.73 versus 1.03 on restricted_mlm, with lower scores better. These trajectories suggest model-dependent effects, not a universal improvement in extension quality.
- Appendix G documents single-seed comparisons for several tasks, recovery resumes, and a budget-free final score-only reevaluation for rust_codecontests. The plotted cost axes also extend beyond $50, so the nominal limits should not be assumed to describe every displayed trajectory without qualification.
Best-so-far trajectories show that earlier useful moves do not guarantee the best final outcome. ARA wins three of five Sonnet 4.6 comparisons, while the paper agent overtakes on triton_cumsum and restricted_mlm. Score direction differs by task, and the cost panels extend beyond the nominal $50 limit, making the appendix’s resume and scoring conditions important to interpretation.
7.5. ARA-Native Review Systems
Review evaluation measures the three automated Seal levels, not human assessment of novelty or significance. All 30 artifacts require validation feedback and pass Level 1 within three iterations; the Level 2 mutation benchmark seeds five defects in each of 23 PaperBench ARAs and scores detection against a hidden injection manifest.
- Table 4 reports 95/115 detected defects (82.6%): 23/23 fabricated claims, 23/23 rebutted-branch leaks, 23/23 over-claims, 21/23 missing falsifications, and only 5/23 orphan experiments.
- The orphan blind spot is consistent with claim-centric traversal missing experiments whose Verifies target does not resolve. Several injections violate checks already specified for Level 1, so the benchmark does not exclusively test higher-order argumentative rigor.
- In 17/23 artifacts, the reported overall mean is rounded upward enough to cross the Accept threshold. Critical findings also fail to produce prescribed dimension-score reductions, supporting the authors’ recommendation to compute verdicts deterministically from findings rather than generate grades independently.
- The paper reuses the reported reproduction success of 64.4% as the Level 3 signal. This is not an independent validation of secure execution or evidence isolation, and passing Level 1 does not by itself prove information completeness.
Detection is complete for three injection classes but drops to 5/23 for orphan experiments, producing 82.6% overall recall. The class-specific asymmetry exposes the limits of claim-centered LLM inspection and supports deterministic reference enumeration; several injected defects are structural violations rather than purely argumentative errors.
8. Related Work
Related work covers machine-readable science, reproducibility infrastructure, negative knowledge, and agent-oriented tooling. Table 5 positions ARA’s contribution as explicit cross-layer bindings rather than replacement of PDFs, repositories, or experiment trackers.
- FAIR, PROV-O, nanopublications, knowledge graphs, RO-Crate, and Whole Tale address metadata, provenance, structured claims, packaging, or environments. The paper argues that these approaches do not jointly provide its execution specifications and decision history.
- Workflow engines and notebooks capture computation without necessarily encoding claim semantics; raw trajectory archives preserve activity without necessarily exposing actionable failure modes.
- ARA differs from post-hoc paper-to-code pipelines by proposing capture of claims, evidence, heuristics, and their bindings at authoring time. Table 5 is a conceptual coverage comparison, not a measured head-to-head evaluation.
The coverage matrix identifies explicit cross-layer bindings as the dimension missing from the other listed artifact types. Its checkmarks express the authors’ conceptual taxonomy; they do not establish that every repository or tracker lacks comparable integrations or quantify an empirical performance advantage.
9. Future Work
Future work proceeds from artifact lineage and maintenance to corpus-level collaboration and cross-disciplinary adaptation. Proposed structured parent–child diffs would reduce construction and re-verification work, while consuming agents could repair stale dependencies and propagate corrections upstream.
- A corpus knowledge graph could align claims, compare cited baselines, expose conflicting trajectories, and support continuous confidence updates from replications and counter-evidence.
- Generalization to wet-lab research requires adapting execution and exploration layers to physical experiments. Section 7.4 also suggests selective trace disclosure and model-class provenance to reduce anchoring by stale heuristics.
10. Limitations
The stated limitations concern evaluation scope, source-bounded fidelity, and deployment prerequisites. Experiments cover ML artifacts; compilation inherits omissions from its available sources, and high-fidelity live capture assumes that a coding agent participates throughout the project.
- The benchmark was constructed by annotators familiar with ARA and the selected papers, leaving performance on unfamiliar or niche-domain artifacts uncertain. Physical experiments and proof-based contributions remain empirically untested.
- Sandboxed execution, content-level anomaly detection, and granular Exploration Graph access control are not implemented. Security and privacy properties in the proposed review architecture are therefore not current guarantees.
- PAPER.md includes an ara_schema version tag, and validators are intended to tolerate unknown fields and missing optional fields. Major-version migration, archival rewriting, long-term checker availability, and deprecation policy remain unresolved.
11. Conclusion
The conclusion presents ARA as infrastructure for publishing, verifying, and extending agent-operable research knowledge. The experiments support improved retrieval and bounded reproduction with enriched structured artifacts, while mixed extension outcomes qualify the broader claim that more complete process documentation always improves downstream discovery.
Appendix
- A. ARA Protocol and Design Rationale derives ten reproduction-critical information categories from a five-paper, 3,050-leaf subset. Combinatorial experiment matrices account for 24.1%, evaluation protocols 18.5%, and hyperparameters 17.2%; the remaining categories cover logging, interpretation, architecture, formulations, implementation tricks, data pipelines, and infrastructure. The key findings are that hyperparameters cover only a minority of requirements, experiment cross-products need explicit enumeration, and high-weight interpretation requirements demand verification of qualitative claims rather than merely runnable code. A.2 distinguishes src_mode: kernel | repo. A.3 demonstrates the paper’s own artifact: PAPER.md indexes the layers; logic/claims.md records status, provenance, falsification criteria, and proof pointers; logic/problem.md links observations to gaps and an insight; logic/solution/heuristics.md records decisions with rationale and sensitivity; trace/exploration_tree.yaml preserves decisions, failures, and experiments; and trace/sessions/ records events, changed files, affected claims, and open threads. The excerpts report different snapshot counts—114 versus 94 exploration nodes, 23 versus 18 heuristics, and 38 versus 36 sessions. The manifest’s NeurIPS 2026 entry has status: draft and does not establish publication or acceptance.
- B. Compiler Skill Details describes an approximately 482-line natural-language skill covering workflow, tool conventions, the directory schema, four compilation stages, and output invariants. The host agent supplies execution mechanics; the skill specifies domain requirements. Required claim fields include Statement, Status, Falsification criteria, and Proof, with the normative specification requiring experiment-ID references rather than file paths. Experiment plans contain directional expectations but place exact results exclusively in /evidence/. The specification requires 15 mandatory files and an exploration tree with at least eight nodes, including a dead_end and a decision, and prohibits invented claims, results, and heuristics. These are implementation requirements, not additional evidence that compilation always satisfies them.
- C. Live Research Manager Details expands the rationale for portable integration, epistemic provenance, and capture-time cross-layer bindings. Closure-driven promotion uses topic abandonment after a default k = 5 subsequent turns with no unresolved question, verbal affirmation, empirical resolution, or downstream artifact commitment. Refuted observations can become dead_end entries. Contradictions preserve both entries and create an unresolved decision node for researcher adjudication. Cross-session continuity uses trace/pm_reasoning_log.yaml and compressed key-context records to retain organizational rationale and conversational context without requiring complete prior transcripts.
- D. Test Corpus lists all 23 PaperBench papers and identifies the 15 included in reproduction. The other eight remain in understanding but are excluded from reproduction because of compute or infrastructure constraints, including multi-day fine-tuning, specialized simulation, full ImageNet sweeps, or large benchmark suites. The harness relaxes PaperBench’s original author-code blacklist so the conventional baseline can use companion repositories. The corpus spans efficiency, alignment, interpretability, RL, scientific ML, generative modeling, optimization, retrieval, evaluation, and adaptation. Although several passages describe all papers as ICML 2024 papers, the selection criteria explicitly include two NeurIPS 2024 workshop development papers.
- E. Understanding Evaluation details question templates, fresh-context evaluation, ternary grading, PDF-gap labeling, exploration-cost accounting, category mechanisms, and statistics. Gap labels are produced by an LLM judge required to cite supporting PDF passages; 64% of requirements receive high-confidence labels. Table 9’s gap counts exceed both the partial/absent requirement count and the 8,921-requirement corpus, so they should not be interpreted as mutually exclusive requirement counts without further clarification. The 24,008-run exploration analysis spans 228 tasks and $63,483 total cost, distinguishing necessary exploration from avoidable downstream rediscovery. Category A emphasizes indexed retrieval, Category B centralized configurations, and Category C access to otherwise absent failure knowledge. Appendix E reports McNemar χ² = 95.15 and p < 10⁻¹⁰, but its difficulty counts of 74, 193, and 172 plus 26 unanswerable questions do not reconcile with 450 unless categories overlap in an unstated way.
- F. Reproduction Evaluation defines the primary metric as sum_i(s_i · w_d_i) / sum_i(m_i · w_d_i), where w_d is 1, 2, or 3 for easy, medium, or hard tasks, s_i sums satisfied requirement weights plus 0.5 times partially satisfied weights, and m_i is the maximum score. It reports Wilcoxon p = 0.028 and exact-binomial p = 0.039 for a summarized 8–2 sign pattern, but the stated binomial probability is not consistent with that pattern under a conventional exact test. Figure 13 and Table 11 provide difficulty-specific results; Table 11’s weighted columns contain nine positive and six negative differences, with no exact ties, while the paper reports 8/5/2 without specifying a tie threshold. Several narrative deltas also disagree with the table, including the stated ftrl advantage despite lower tabulated ARA performance. The aggregate rows are not transparently reconciled with the listed per-paper values, and rice difficulty values are explicitly interpolated. The reported overall reproduction rates should therefore not be presented as independently reconstructed from the appendix.
- G. Extension Evaluation documents task selection, construction, baselines, harness engineering, score extraction, and five case studies. optimize_llm_foundry lacks a published MALT corpus, while small_scaling_law has sparse, strategically uninformative traces. Construction uses per-run extraction, provenance-tagged merging, and a beat-reference filter enforced during extraction and merge; the synthesized paper baseline deliberately preserves successful methods rather than rejected alternatives. Harness repairs include larger SDK buffers, blocking mass-batch scoring, stop-loop pushback and recovery resumes, and task-specific scorer timeouts. Scores come from canonical scorer outputs, not agent commentary, and costs are reconstructed from usage records and reconciled with SDK totals. The triton_cumsum case contrasts useful Sonnet 4.5 autotune guidance with Sonnet 4.6 anchoring to a trace-recommended redesign. rust_codecontests shows earlier adoption of a hand-coded Rust library, but includes a recovery resume and budget-free final scoring. nanogpt_chat_rl shows trace-guided continuation of the reference algorithm and a documented degenerate-output filter rather than repeated algorithm redesign. fix_embedding shows comparable execution of the published adapter recipe followed by different post-reference searches, including repeated permutation-recovery attempts only in the paper arm. restricted_mlm shows focused tuning under Sonnet 4.5 but fragmented exploration under Sonnet 4.6; Figure 15 also reconstructs the ARA-4.5 starting anchor after its initial baseline measurement crashed. Figures 14–15 provide the Sonnet 4.5 comparisons. Code and reconstructible event logs are located under code/extension-harness/, code/rebench-pipeline/, code/artifacts/rebench-<task>/, and each run’s trace.jsonl.
- H. Review System Evaluation specifies Level 1 directory, file, schema, minimum-count, and bidirectional reference checks, including experiment Verifies targets. It reports first-pass success of 0/30, with all artifacts passing within three iterations; compiler failures comprise dangling references (42%), missing fields (31%), insufficient exploration-node counts (14%), parse errors (8%), and missing files (5%). The Level 2 mutation benchmark matches findings to hidden injected entities without using generated grades as ground truth; Table 14 localizes the orphan-experiment misses and the missing-falsification misses in bam and bbox. Critical rebutted-branch findings do not produce the prescribed exploration-integrity score reductions, illustrating finding–score decoupling. The appendix’s suggestion that all residual fidelity errors must reflect absent source information does not follow from structural validation alone. Level 3 reuses the reproduction evaluation rather than introducing an independent execution benchmark.
Brief Thoughts
The clearest contribution is the explicit claim–experiment–code–evidence graph, coupled with structured negative knowledge. The retrieval results support making operational information accessible to agents, but additional rubric and transcript inputs prevent attributing all gains to the directory format. An information-matched baseline is needed to separate organization from enrichment, and the reproduction reporting needs reconciliation before its exact aggregate advantage can be treated as independently verified.
The mixed extension outcomes qualify the introduction’s claim of consistent superiority: preserved history can shorten rediscovery while also restricting or dispersing search. The auditor results support deterministic reference checks and verdict calculation alongside LLM-generated findings. Small trajectory samples, inconsistent appendix summaries, evidence access during reproduction, and unimplemented security controls limit ARA’s demonstrated status to a research-infrastructure prototype rather than a validated production publication system.