AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation | Summary
27 May 2026 | Paper Review AI Research Agents Multi-Agent Coordination Agent's MemoryContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 AUTOSCIENTISTS: Long-Running Self-Organizing Agent Teams
- 4 Experiments
- 5 Limitations and Future Work
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation.
- 2026-05-27 (arXiv)
- Gao, Shanghua, Fang, Ada, Zitnik, Marinka.
- Harvard University
- Paper
- Project Page
Summary
- AUTOSCIENTISTS coordinates long-running computational experimentation without a central planner agent. Agents critique proposals, form and reorganize teams, and share a champion program, experimental results, and failed directions as evidence accumulates.
- On 24 BioML-Bench tasks, AUTOSCIENTISTS achieves a mean leaderboard percentile of 74.40 (6.20)% versus 66.07 (7.38)% for Autoresearch, a gain of 8.33 percentile points. AUTOSCIENTISTS, Autoresearch, and rerun Biomni receive matched experimental hardware and time budgets; other published baselines use different hardware protocols for non-imaging tasks.
- For GPT nanochat training optimization, AUTOSCIENTISTS reportedly reaches val_bpb ≈ 0.978 in 34 experiments versus 65 for Autoresearch; this measures experiment efficiency, not a 1.9× wall-clock speedup. Starting from the same stronger champion and dead-end registry, it accepts seven updates over 93 experiments and reaches 0.9730, whereas Autoresearch accepts none over 100 experiments.
- A Kermut extension developed on ACE2–Spike binding transfers unchanged to 217 ProteinGym assays, increasing average Spearman’s ρ from 0.657 to 0.700 while slightly worsening MSE from 0.605 to 0.611. The results support improved experimental search and ranking performance, but do not establish lower total cost, clinical validity, or statistically reliable benefits for every coordination component.
1 Introduction
The introduction identifies a mismatch between fixed research workflows and long-running experimentation: competing hypotheses emerge, promising directions become exhausted, and failures must inform later decisions. Present Work. introduces shared-state coordination without a central orchestrator agent, allowing teams to redirect effort as evidence changes.
- The contribution concerns computational experimentation across biomedical ML, language-model training, and protein fitness prediction; autonomous wet-lab research is not evaluated.
- The website is https://autoscientists.openscientist.ai and the code repository is https://github.com/mims-harvard/AutoScientists.
The diagram connects team formation and reformation to a common champion, experiment log, forum, queues, and dead-end records. Stagnation prompts renewed discussion, but the parallel-experiment blocks describe system capacity rather than demonstrating concurrent GPU training in the one-GPU benchmark.
2 Related Work
Related work distinguishes scientific-agent capabilities from the organization of multi-agent collaboration. AUTOSCIENTISTS focuses on how agents select and sustain research directions, rather than introducing a new base language model or scientific tool.
AI Agents for Scientific Research.
AI Agents for Scientific Research. reviews literature synthesis, hypothesis generation, biomedical analysis, code execution, and iterative algorithm development. AUTOSCIENTISTS differs from fixed PI–scientist–critic or manager-led pipelines by letting agents determine research directions through shared discussion.
- Discussion critiques and filters proposals before experimental compute is spent, while retaining multiple research directions rather than requiring agreement on one hypothesis.
Coordination of Multi-Agent Systems.
Coordination of Multi-Agent Systems. emphasizes that adding agents or communication does not automatically improve performance. Findings on collaboration structure, diversity, coordination costs, and context management motivate the single-agent comparisons and coordination ablations.
3 AUTOSCIENTISTS: Long-Running Self-Organizing Agent Teams
AUTOSCIENTISTS formulates long-running experimentation as iterative program search with persistent agent identities and accumulated knowledge. The method combines dynamically formed research teams, role-specific proposal and execution loops, and shared evidence carried between invocations.
3.1 Problem Formulation
The input comprises a task description, an optional initial program p₀, a dataset D, and an evaluation metric ℓ. The objective is p* = arg max over p ∈ P of ℓ_eval(p; D), where P is the explored program space and the metric is oriented so that larger values are better.
- Development feedback comes from a validation set or cross-validation over the training set.
- Final performance uses a held-out test metric when available; otherwise, the development evaluation protocol supplies the reported result.
3.2 AUTOSCIENTISTS Approach
Overview. alternates discussion and execution phases. Agents read the evidence, organize around research directions, execute experiments, and reopen discussion when progress stagnates; shared state mediates coordination rather than a planner agent.
- The approach is described as LLM-agnostic, although the reported experiments use one coding-agent backend.
Discussion and Self-Organization. starts without team assignments or a predefined partition of the search space. Agents propose directions, rank hypotheses, critique earlier posts, and consolidate the discussion into a roster mapping teams to research axes and members.
- Discussion ends after a majority vote for [DISCUSS-DONE]; the alphabetically-last participating analyst writes the roster as a deterministic tie-breaker, not as a privileged planner.
- Reformation can create, merge, split, retire, or rebalance teams. The main description requires endorsement from affected teams, and gives no improvement in the last 10 experiments as an example stagnation trigger.
Long-Running Parallel Experiments. divides work between analyst agents and experiment agents. Analysts audit untested directions and queue proposals; experiment agents claim a proposal, modify the shared champion, train and evaluate the candidate, and publish its outcome.
- Each heartbeat reads shared state, performs role-specific work, records results, and persists knowledge for subsequent sessions.
- Improvements inside an empirically estimated noise band require second-seed confirmation before champion promotion.
- Cross-readable dead-end registries and effect-size rankings discourage repeated failures and prioritize underexplored directions.
Shared State. contains the champion and reproduction instructions, a complete experiment log, a structured research forum, and cross-readable team queues, dead-end registries, and hypothesis documents. Output. comprises the final champion, a model card, and a research findings report tracing successful and rejected directions.
- Figure 2 illustrates the hERG model card; Appendix E provides GPT nanochat artifacts.
- Analyst-recorded explanations document the search process but do not independently establish causal mechanisms.
The hERG artifact combines model specification, test evaluation, validation trajectory, seed sensitivity, ablations, and use limitations. It demonstrates experimental documentation beyond a final score, but is not independent evidence of clinical validity.
4 Experiments
The experiments evaluate biomedical ML pipeline development from task specifications, GPT training-code optimization under short training budgets, and extension of a strong protein fitness predictor. Component ablations examine coordination mechanisms on selected tasks.
4.1 Implementation Details
All agents use Claude Code with Claude Sonnet 4.6, and Autoresearch uses the same backend. A deterministic monitor repeatedly invokes agents in heartbeat loops; the default crew contains 3 analyst agents and 6 experiment agents.
- H100 GPUs are available for experiments, but the one-GPU-per-task BioML-Bench budget requires GPU-bound experiments to execute sequentially.
- Persistent state connects successive LLM sessions; long-running coordination does not require an uninterrupted model context.
4.2 End-to-End Biomedical Machine Learning with AUTOSCIENTISTS
Setup. evaluates 24 BioML-Bench tasks: 4 biomedical imaging, 9 drug discovery, 6 protein engineering, and 5 single-cell omics tasks. Agents receive task descriptions, training data, test inputs, submission examples, and validation feedback; hidden test labels and private graders remain outside their workspace.
- AUTOSCIENTISTS discussion prompts additionally include an LLM-generated model-paradigm menu for each task family to encourage diverse proposals.
- AUTOSCIENTISTS, Autoresearch, and rerun Biomni receive 1 H100 GPU, 16 CPUs, and 48 GB memory for 4 hours per non-imaging task and 16 hours per imaging task. Published non-imaging baselines instead use an 8-hour CPU-only protocol.
- Outcomes are leaderboard percentile, above-median submissions, medal attainment, and completion. The evaluated agents do not use web search or fetch.
Results. show 74.40 (6.20)% mean leaderboard percentile and completion of all 24 tasks. The largest domain gain over Autoresearch is in drug discovery, from 46.16 (10.59)% to 64.52 (8.37)%; imaging reaches 45.75 (22.18)%, and single-cell omics reaches 88.00 (9.70)%.
- Protein engineering is near the percentile ceiling: AUTOSCIENTISTS and Autoresearch both obtain 96.97 (3.03)%, although their mean ranks differ, 1.50 versus 2.00.
- Figure 3 reports 75.0% above the public leaderboard median and 70.8% receiving any medal.
- Aggregate superiority does not mean a win on every task metric; Table S7 includes tasks won by Biomni or Autoresearch.
Domain averages show the largest percentile advantage over Autoresearch in drug discovery and a tied percentile ceiling in protein engineering. Protein-engineering mean ranks still differ, showing that coarse leaderboard percentiles can conceal differences in task scores; large imaging standard errors also limit domain-wide conclusions.
AUTOSCIENTISTS leads the plotted aggregate outcomes with 74.4% mean leaderboard percentile, 75.0% above-median submissions, and 70.8% medal attainment. The error bars summarize variation across tasks rather than repeated system runs, and several published baselines use different non-imaging hardware budgets.
Inspection of shared-state records provides qualitative evidence that deliberation changes experiment selection. Figure 5 shows proposal diversification, recognition of saturated directions, cross-team hypothesis transfer, and retirement of a dead end after stagnation.
- These selected traces illustrate the proposed mechanism but do not quantify its frequency or isolate the causal effect of individual interactions.
4.3 GPT Training Optimization with AUTOSCIENTISTS
Setup. assigns each GPT nanochat experiment a 5-minute training run on one H100 GPU and evaluates validation bits-per-byte, val_bpb, with lower values better. Both systems use the same repository and backend, and progress is compared by experiment count rather than wall-clock time.
- The baseline-start regime begins at val_bpb = 0.998.
- The champion-start regime begins at val_bpb = 0.9777 after 50 prior AUTOSCIENTISTS experiments; both systems receive the same EXPLORED.md record of dead-ended directions.
From Autoresearch Baseline. reports that AUTOSCIENTISTS reaches val_bpb ≈ 0.978 in 34 experiments, compared with 65 for Autoresearch. At the matched 50-experiment endpoint in Table S2, the values are 0.9777 and 0.9790, respectively.
- The run forms architecture, schedule, and optimizer teams, maintaining several research directions.
- The 1.9× comparison concerns experiments needed to reach a target loss. It does not demonstrate a hardware-parallel wall-clock speedup, and Figure 4a itself displays only the first 50 experiments.
From an AUTOSCIENTISTS Champion. reports seven accepted updates over 93 experiments, reaching val_bpb = 0.9730. Autoresearch accepts no updates over 100 experiments; its best attempted candidate reaches 0.9783, worse than the starting champion.
- Accepted directions include query-key normalization order, matrix initialization, value-embedding gate width, final-learning-rate fraction, softcap value, compile autotuning, and noise-floor recalibration.
- Query-key normalization order is never proposed in the 100 Autoresearch attempts, giving a concrete example of differing hypothesis coverage.
- One accepted update recalibrates the starting champion rather than changing the model. Appendix B describes a 100-experiment budget, whereas the main text reports 93 completed AUTOSCIENTISTS experiments.
The baseline-start panel shows lower best-so-far loss for AUTOSCIENTISTS later in the displayed 50-experiment window. The champion-start panel shows further accepted gains against an unchanged Autoresearch champion despite equal access to prior negative knowledge. These experiment-count trajectories do not measure wall-clock parallel speedup; the caption’s 65-experiment baseline comparison extends beyond the displayed panel.
4.4 Extending Kermut on ProteinGym Supervised Fitness Prediction
Setup. starts from Kermut, a Gaussian-process protein variant-effect predictor, and develops modifications on ACE2–Spike binding. Agents optimize mean out-of-fold Spearman correlation over random, modulo, and contiguous split schemes for 10 cycles on one H100 GPU, without human intervention.
- A cycle corresponds to each experiment agent completing one queued experiment.
- The discovered AUTOSCIENTISTS-Kermut recipe is frozen before application to all 217 ProteinGym supervised substitution assays.
Results. improve development-assay Spearman’s ρ from 0.747 to 0.840, a 12.5% relative gain. The predictor combines three GPs, Kermut’s structure kernel, expanded zero-shot features, and quantile-warped targets; the main text also describes greedy diversity-based feature selection.
- Across 217 assays, average Spearman’s ρ increases from 0.657 (0.008) to 0.700 (0.007), an absolute gain of 0.043 and a 6.5% relative improvement.
- Average correlation improves under contiguous, modulo, and random CV, but average MSE increases from 0.605 to 0.611.
- Frozen cross-assay application supports reuse beyond the development assay. ProteinGym evaluates out-of-fold predictions rather than a separate hidden test set, and the 217-assay aggregate includes the development assay.
The frozen predictor increases average Spearman correlation under contiguous, modulo, and random CV across 217 assays, with overall ρ rising from 0.657 to 0.700. Average MSE rises from 0.605 to 0.611, showing that better variant ordering need not improve regression accuracy.
4.5 Ablations of AUTOSCIENTISTS
Setup. removes analyst roles, cross-agent feedback, self-organization, or both communication and shared state on TDC-hERG, Cell-Cell Communication, Human Plasma-Protein Binding, and GPT training optimization. The main text presents these as controlled comparisons with the same backend, task interface, starting program, and experimental budget.
- No analyst transfers proposal generation and knowledge maintenance to experiment agents.
- No cross-agent feedback removes proposal and result discussion; Appendix D also disables stagnation-triggered re-discussion.
- Independent agents keep private champions and knowledge, preventing inheritance or combination of another agent’s results.
Results. favor the full system on every primary metric in Table 3, with task-dependent sensitivity. Removing analysts reduces hERG AUROC from 0.867 to 0.738; removing feedback reduces plasma-protein-binding Pearson correlation from 0.8729 to 0.7144; removing self-organization raises GPT val_bpb from 0.9777 to 0.9833.
- Independent agents reduce the Cell-Cell Communication score from 0.924 to 0.435 and also reach 0.9833 on GPT.
- The patterns are consistent with complementary coordination functions, but Appendix D reports one GPT run per condition and unequal completed experiment counts.
- Disabling re-discussion in the no-feedback condition prevents attribution of its entire effect to communication alone.
The full system has the best primary score on all four tasks, but the largest loss comes from different removals on different tasks. The pattern is consistent with complementary functions, while single-run conditions, unequal appendix experiment counts, and disabled re-discussion in the no-feedback condition limit causal attribution.
Four selected interaction records illustrate proposal diversification, saturation analysis, cross-team hypothesis transfer, and team merging after an unproductive Personalized PageRank direction. They connect the coordination design to recorded search decisions without establishing the frequency or average benefit of these behaviors.
5 Limitations and Future Work
AUTOSCIENTISTS is designed to improve search under fixed experimental compute, not to minimize LLM calls or total monetary cost. Table S8 estimates BioML-Bench LLM costs of $3984.7 for AUTOSCIENTISTS versus $910.1 for Autoresearch, excluding GPU compute.
- One-GPU evaluation does not fully test concurrent GPU experimentation; scaling under larger GPU budgets remains future work.
- Agent count is fixed before a run. Preliminary sweeps show task-dependent best crew sizes and degradation at 14 agents.
- Additional evaluation limitations are incomplete multi-run replication, single-run component ablations, and testing only one base LLM.
6 Conclusion
The conclusion links sustained search to proposal critique, shared successful and failed experiments, and team reorganization as productive directions shift. The comparisons support improved trial selection and continued improvement from a strong starting program more directly than universal superiority or lower end-to-end cost.
Appendix
- Reproducibility Statement provides launch scripts and public assets at https://github.com/mims-harvard/AutoScientists. Official repositories include https://github.com/science-machine/biomlbench, https://github.com/OATML-Markslab/ProteinGym, https://github.com/karpathy/autoresearch, https://github.com/snap-stanford/biomni, and https://github.com/petergroth/kermut; Table S1 records versions, commits, and licenses. Runs use successive Claude Code sessions without human intervention during the specified budget. Impact Statement warns about over-trusting discovered models, benchmark overfitting, and propagation of erroneous hypotheses; biomedical outputs require expert review and independent validation.
- A Implementation Details and Algorithmic Protocols specifies launch, team formation, and normal execution. A.1 Setup and Launch initializes shared state and unassigned agents, then posts a discussion trigger; A.2 Heartbeat Protocol dispatches discussion, no-team exit, pending-result recovery, or role-specific work. In A.3 Self-Organized Team Formation, agents propose and rank directions, critique earlier contributions, and cast [DISCUSS-MORE] or [DISCUSS-DONE] votes; a majority of DONE votes permits the alphabetically-last participating analyst to write the roster. Reformation may create, merge, split, retire, or rebalance teams, with later sessions adopting the persisted roster. A.4 File Discovery uses list–decide–read access, and A.5 Queue and Claim Protocol uses version-token optimistic locking to prevent duplicate experiment claims.
- A.6 Noise-Aware Champion Validation accepts an improvement directly when Δ > Mσ, with M = 2; when 0 < Δ ≤ Mσ, a second seed must also improve over the champion, and Δ ≤ 0 is rejected. The noise estimate is σ = sqrt[(1/(2n)) Σᵢ(ℓ₁,ᵢ − ℓ₂,ᵢ)²], begins after at least three paired runs, and is locked after five pairs; before estimation, the system uses a conservative default noise band. This protects the shared baseline against champion pollution, but is a heuristic promotion rule rather than a multiple-comparison-controlled significance test.
- A.7 Analyst Proposal Protocol generates two proposals per invocation using numeric-parameter coverage audits and empirical mean absolute effect sizes. Directions with fewer than three trials receive an exploration bonus, while effects below the noise floor are deprioritized. At least one proposal must satisfy a bold-move criterion—such as a ≥10% parameter-count change, a confirmed bug fix, an axis repeatedly flagged as untested, or a hypothesis-tension probe—or receive a public [EXEMPT] justification. The proposals must target different directions, avoid repeated same-direction runs and rejected ranges without new justification, and include a mechanism-based follow-up after a champion update.
- B Extended Ablation Results includes B.1 Comparison Against Autoresearch at Matched and Extended Compute: Table S2 reports 0.9777 versus 0.9790 after 50 baseline-start experiments and seven versus zero accepted updates from the shared champion. B.2 Team-Size Sensitivity: Parallelism Gain and Oversubscription compares crews of 2, 4, 9, and 14 working agents, with experiment/analyst compositions of 1+1, 2+2, 6+3, and 9+5. The best crew varies by task, and 14 agents underperform each task’s best configuration. Smaller productive crews reach similar GPT losses with different experiment counts. The approximately 3.25× speedup for 9 versus 2 agents is projected from heartbeat counts under assumed fully parallel execution; the actual agents were invoked iteratively rather than fully in parallel.
- C Run-to-Run Stability reports three independent GPT cold starts ending at val_bpb values of 0.9777, 0.9795, and 0.9780 after 75, 64, and 62 experiments. Their reported mean is 0.9784 with sample standard deviation 0.0010, and acceptance rates range from 9.7% to 10.9%. Different intermediate directions reach a similar performance region, providing limited replication evidence for the full system without establishing comparative variance against Autoresearch.
- D Per-Experiment Trajectories for Ablation Runs covers D.1 No Self-Organization (abl-no-self-org), D.2 No Cross-Agent Communication (abl-no-cross-agent), D.3 No Analyst Role (abl-no-analyst), and D.4 Independent Agents (abl-independent). Their final GPT val_bpb values are 0.9833, 0.9814, 0.9817, and 0.9833 over 47, 50, 50, and 50 experiments, compared with 0.9777 over 71 for the full-system reference. The no-self-organization run was stopped after a budget revision; the no-communication condition also disables re-discussion; and the no-analyst run still produces cross-team improvements through shared artifacts and suggestions. Independent agents repeatedly rediscover the batch-size reduction and cannot combine divergent champions. The appendix reference Autoresearch run eventually reaches 0.9773 after 83 experiments, slightly below the full-system endpoint but with more trials; these single-run, unequal-length comparisons do not establish a statistically reliable component ranking.
- E AUTOSCIENTISTS Output on GPT nanochat presents E.1 Model Card and E.2 Research Insights for a 75-experiment cold-start run reaching val_bpb = 0.97769. The model card describes an English, 87 M-parameter decoder-only transformer trained on a FineWeb shard with the upstream BPE-8192 tokenizer; it is a short-budget research artifact, not a production language model or a validated downstream-task system, and inherits source-data biases. Training uses DEPTH = 7, ASPECT_RATIO = 96, a hybrid Muon/AdamW optimizer, bf16, and torch.compile; 300 s on one H100 80 GB GPU yields 1,293 optimizer steps, 339 M processed tokens, and 58 GB peak VRAM. Evaluation uses a pinned 40 M-token held-out FineWeb shard and reports reproducibility within an empirical noise floor. Research Insights records approximately 6.4 hours of search and 8 KEEPs, comprising a baseline lock and 7 improvements: TOTAL_BATCH_SIZE 2¹⁹→2¹⁸, width growth, fewer Newton–Schulz iterations, two attention-window reductions, exact Polar-Express normalization, and DEPTH 8→7. Analyst notes attribute these changes to throughput, capacity coupled with learning-rate scaling, and optimizer quality; schedule-shape changes and other rejected probes delimit the productive region, while a failed second-seed confirmation prevents promotion of short_window_ratio = 1/16. Team records show cross-team reuse of earlier improvements, but causal explanations, ordering effects, and generalization remain limited by a single task, backend, run, and 300 s training regime.
- F Implementation Details of BioML-Bench defines compute resources and iterative validation leaderboarding in F.1 Setup. AUTOSCIENTISTS, Autoresearch, and rerun Biomni use 4-hour H100-backed budgets for non-imaging tasks and 16 hours for imaging, whereas published non-imaging baselines use 8-hour CPU-only budgets. F.2 Task Datasets and Evaluation Metrics specifies domain-dependent splits and scoring. F.2.1 Biomedical Imaging covers kaggle-histopathologic-cancer-detection with ROC-AUC and a random training split; kaggle-osic-pulmonary-fibrosis-progression with modified Laplace log-likelihood and patient-level CV; kaggle-rsna-miccai-brain-tumor-radiogenomic-classification with ROC-AUC and patient-level CV; and kaggle-uw-madison-gi-tract-image-segmentation with 0.4 Dice + 0.6 (1 − HD3D) and held-out patients and treatment dates.
- F.2.2 Drug Discovery uses five-fold Murcko-scaffold CV and label-blinded held-out tests. In source order, tdcommons-bbb-martins evaluates BBB permeability with ROC-AUC; tdcommons-caco2-wang evaluates cell permeability with MAE; tdcommons-cyp2d6-substrate-carbonmangels evaluates substrate activity with PR-AUC; tdcommons-herg evaluates channel inhibition with ROC-AUC; and tdcommons-lipophilicity-astrazeneca evaluates lipophilicity with MAE. polaris-adme-fang-hclint-1, polaris-adme-fang-hppb-1, and polaris-adme-fang-solu-1 evaluate clearance, plasma-protein binding, and solubility with Pearson correlation. polaris-pkis2-egfr-wt-c-1 evaluates imbalanced EGFR inhibition with PR-AUC.
- F.2.3 Single Cell Omics evaluates open-problems-predict-modality with RMSE on held-out site/donor batches and open-problems-single-cell-perturbations with MRRMSE, using a 20% within-cell-type validation split. open-problems-cell-cell-communication-ligand-target has 81 labelled pairs and 731 unlabelled test pairs, uses a random 80/20 validation split, and reports bounded score = 1 − 1/(1 + OR/2); the main tables’ Odds Ratio label is therefore shorthand for a transformed score. open-problems-label-projection uses weighted F1 on a held-out batch. open-problems-spatially-variable-genes uses Kendall’s τ for 210 cortex genes, with cerebellum labels available to check an otherwise unsupervised spatial statistic.
- F.2.4 Protein Engineering evaluates out-of-fold predictions without a separate held-out test set. Substitution tasks use fold_random_5, fold_modulo_5, and fold_contiguous_5 for SPIKE_SARS2_Starr_2020_binding, SBI_STAAM_Tsuboyama_2023_2JVG, CBX4_HUMAN_Tsuboyama_2023_2K28, and PSAE_PICP2_Tsuboyama_2023_1PSE; indel tasks Q8EG35_SHEON_Campbell_2022 and CSN4_MOUSE_Tsuboyama_2023_1UFM use fold_random_5. The final metric is mean Spearman correlation in raw target space.
- F.3 Performance Across Independent Runs reports hERG AUROCs of 0.867, 0.830, and 0.862, with mean 0.853 and standard deviation 0.020; its assertion that all three beat the Table S7 baselines conflicts with Biomni’s 0.85464. F.4 Performance Across Wall-Clock Time tracks six protein-engineering tasks over 4-hour budgets, where development CV and final scoring use the same protocol. F.5 Behavior of AUTOSCIENTISTS on BioML-Bench Tasks classifies pipelines into seven non-exclusive method categories. Imaging fine-tunes pretrained CNNs on three tasks and combines frozen CT features with regression and nearest neighbors on OSIC; drug discovery uses fingerprint-based boosting on six tasks, alongside neural, kernel, and stacking components. Single-cell omics uses two training-free heuristics, two linear-model pipelines, and one residual MLP; protein engineering uses frozen ESM-2 features on five tasks and LoRA fine-tuning on two. Overall, Linear/Ridge occurs in 10/24 tasks, boosting in 8/24, frozen foundation-model features in 7/24, and foundation-model fine-tuning in 6/24, usually as components of broader pipelines rather than mutually exclusive strategies.
- G AUTOSCIENTISTS-Kermut for ProteinGym ACE2-Spike binding DMS and G.1 Setup specify a frozen three-GP recipe. Inputs include 1280-dimensional ESM-2 embeddings, ProteinMPNN log-probabilities, ESM-2 scores, 15 additional zero-shot predictors, and four contact-map features; standardization, missing-value imputation, and quantile target normalization are fitted within each training fold. The Kermut structure kernel combines ProteinMPNN distribution and scalar log-probability distances with Cα distances, then mixes with each member’s sequence kernel: Matérn-3/2 ARD on 1303 features, Matérn-5/2 ARD on 1280-dimensional embeddings, or a linear kernel on 18 scalar features. Region-aware fixed noise varies with distance between training and held-out fold indices, while additional homoscedastic noise is learned. Each member uses 1000 AdamW steps with learning rate decaying from 10⁻¹ to 10⁻³ and seed 2024; preprocessing and models are refitted for each prescribed CV fold. Converged predictions are averaged, numerically failing members are dropped, and MSE is computed against z-scored targets with predictions clipped to [−10, 10].
- G.2 Ablations of AUTOSCIENTISTS-Kermut reports average development-assay Spearman’s ρ of 0.8407 for the full predictor, 0.8231 without ensembling, 0.8070 without quantile normalization, and 0.8174 without extra zero-shot scores. Table S9 therefore identifies removal of quantile normalization—not ensembling—as the largest average ranking loss, contrary to the surrounding prose. Removing quantile normalization improves MSE from 0.4527 to 0.3501, reinforcing the distinction between rank optimization and regression-scale accuracy. Removing the extra zero-shot scores worsens both average correlation and MSE, supporting their contribution in this assay without establishing reliability across repeated discovery runs.
Brief Thoughts
The clearest contribution is an operational design for accumulating and reusing experimental evidence: a common champion, rejected-direction memory, critique before execution, and explicit reorganization rules. Continued GPT improvement from a shared strong starting point and frozen cross-assay application of AUTOSCIENTISTS-Kermut provide the strongest evidence that the system can do more than repeat isolated local searches.
The results demonstrate experimental-search gains rather than overall resource efficiency or measured large-scale parallel speedup. Higher LLM spending, limited replication, coupled ablations, unequal appendix run lengths, and internal reporting inconsistencies constrain causal claims about individual components; replicated comparisons matching both experimental and LLM budgets would better test those claims.