Gorio Tech Blog search

Towards Autonomous Mathematics Research | Summary

|

Contents

This article explains the key points of Towards Autonomous Mathematics Research.

  • 2026-02-10 (arXiv)
  • Feng, Tony, Trinh, Trieu H., Bingham, Garrett, Hwang, Dawsen, Chervonyi, Yuri, Jung, Junehyuk, Lee, Joonkyung, Pagano, Carlo, Kim, Sang-hyun, Pasqualotto, Federico, et al.
  • Google DeepMind, UC Berkeley, Brown University, Yonsei University, Concordia University, Korea Institute for Advanced Study, UC San Diego, Caltech
  • Paper

Read this article in Korean


Summary

  • Towards Autonomous Mathematics Research introduces Aletheia, a Gemini Deep Think–based agent that separates solution generation, natural-language verification, and revision. The central problem is transferring competition-mathematics reasoning to professional research, where specialized literature, long arguments, ambiguous statements, and citation accuracy impose additional reliability requirements.
  • Aletheia generated the mathematical content of Eigenweights for Arithmetic Hirzebruch Proportionality without essential human intervention and contributed to several collaborative papers. Deployment on 700 Erdős problems yielded 13 meaningfully correct responses after human review, including two complete novel resolutions and two novel partial resolutions; on FirstProof, a best-of-2 evaluation judged six of ten problems solved, with P8 receiving only 5/7 favorable expert assessments.
  • These results demonstrate selective research assistance rather than consistently reliable autonomous mathematics: 137 of 200 adjudicated Erdős candidates were fundamentally flawed, and individual FirstProof runs contained substantive errors. The paper proposes separate autonomy and mathematical-significance axes, together with Human-AI Interaction Cards, to distinguish AI capability milestones from mathematical advances.

1. Introduction

Competition problems are generally self-contained, whereas research requires specialized knowledge and synthesis across an extensive literature. Aletheia combines a stronger Gemini Deep Think model, inference-time scaling, and search and browsing; the introduction identifies autonomous eigenweight calculations, collaborative independence-polynomial results, an extensive Erdős evaluation, and improved intermediate propositions as its principal evidence.

  • Table 1 separates the degree of AI involvement from mathematical significance. None of the reported results is classified as a Major Advance or Landmark Breakthrough.
  • The autonomous eigenweight work concerns mathematical content, not human-free publication: humans wrote the final papers and retained responsibility for correctness, exposition, and attribution.
  • Prompts and outputs are available at https://github.com/google-deepmind/superhuman/tree/main/aletheia.

The matrix places eigenweights in the essentially autonomous, publishable-research cell, while independence polynomials and generalized Erdős-1051 occupy the collaboration cell. Empty Major Advance and Landmark Breakthrough rows distinguish these demonstrations from higher-significance mathematical results; the footnote identifies Level 2 works as submitted for publication.

Classification of the reported AI-assisted mathematics by autonomy and mathematical significance.
Classification of the reported AI-assisted mathematics by autonomy and mathematical significance.

2. The Aletheia agent: From Olympiads to Research-level Mathematics

Aletheia orchestrates three subagents—Generator, Verifier, and Reviser—each of which internally calls a Gemini base model. It repeats generation and checking until the Verifier approves a solution or a preset attempt limit is reached, operating entirely in natural language rather than the formal languages used by AlphaGeometry and AlphaProof.

  • Figure 1 distinguishes regeneration after a critically flawed candidate from revision when only minor fixes are needed.
  • The architecture addresses errors and hallucinations that persist despite the model’s broad pretrained knowledge of mathematics.

The Verifier routes critically flawed candidates back to generation and candidates needing minor fixes to the Reviser, while approved candidates become final outputs. This is an iterative natural-language checking workflow, not a formal proof-certification pipeline.

Aletheia's Generator–Verifier–Reviser workflow powered by Gemini Deep Think.
Aletheia's Generator–Verifier–Reviser workflow powered by Gemini Deep Think.

2.1. Scaling Laws and the Evolution of Deep Think

The inference-scaling experiments ran each model exactly once per problem at each compute scale, disabled tools, and used human grading. On the 30-problem IMO-Proof Bench Advanced subset, the January 2026 Deep Think model required approximately 100x less compute to match the July 2025 model’s performance; FutureMath Basic showed benefits from scaling but substantially lower accuracy on Ph.D.-level exercises.

  • Figure 2 shows gains with compute followed by plateaus or fluctuations, as well as higher scores from Aletheia’s harness.
  • The July 2025 model perfectly solved five of six IMO problems. The January 2026 model made further progress on IMO 2025 Problem 6 and IMO 2024 variants, but the authors explicitly caution about possible exposure to public problems.
  • The paper mentions new training improvements without providing enough detail to reconstruct them. Its scaling results are empirical, not an explicit analytic law.

2.2. Developing Agentic Harnesses for Research-Level Math

The harness separates a candidate’s final output from its intermediate thinking tokens and applies verification scaffolding, allowing the model to identify errors overlooked during generation. Aletheia achieved a 93% overall score on IMO-Proof Bench Advanced without tools and 96% conditional accuracy on the 29 of 30 problems it answered; on FutureMath Basic, it answered fewer than 60% of problems with conditional accuracy exceeding 82%.

  • Aletheia outperformed same-base-model Deep Think across the tested compute scales, although its dynamic compute usage could not be precisely controlled.
  • The February 2026 Gemini 3 base model subsequently yielded a 95% score on IMO-Proof Bench Advanced without tools.
  • The authors view abstention as useful when expert verification capacity is limited. Their explanations involving bluffing incentives or misleading thinking context remain hypotheses.

The Olympiad panel shows the January 2026 model outperforming the July 2025 model at comparable compute, while the Ph.D.-exercise panel remains at substantially lower scores. Aletheia's marked results exceed the corresponding Deep Think curves, supporting a harness benefit beyond scaling alone; its dynamic budget limits exact compute matching.

Human-graded performance against inference-time compute on IMO-Proof Bench Advanced and FutureMath Basic.
Human-graded performance against inference-time compute on IMO-Proof Bench Advanced and FutureMath Basic.

2.3. Importance of Tool Use

Search and browsing reduced obvious fabricated citations when combined with extensive tool-use training; internet access alone was insufficient. Remaining errors included misrepresenting results from real papers, while adding Python produced only marginal improvements against computational hallucinations.

  • Figure 3 illustrates a fabricated reference used to justify a claim about the pretzel knot P(−3, 5, 13).
  • Figure 4 illustrates a different failure: Galambos’s paper exists, but it does not contain the asserted classical result about prime factors.
  • Checking a reference’s existence is therefore insufficient; its mathematical content must also support the argument.

The model cites a real Galambos paper but attributes to it a prime-factor statement that cannot be found there. This separates bibliographic authenticity from mathematical support and illustrates a failure mode that persists after tool-use training.

Incorrect attribution of a mathematical result to an existing reference.
Incorrect attribution of a mathematical result to an existing reference.

3. Summary of Mathematical Research Results

The research papers summarized here used Aletheia with a November 2025 base model, intermediate between the July 2025 and January 2026 versions compared earlier. Human authors wrote every final paper from the model outputs because authorship entails accountability for all contents, including attribution and exposition that formal verification alone cannot guarantee.

3.1. Milestone A: Eigenweights for Arithmetic Hirzebruch Proportionality

Eigenweights for Arithmetic Hirzebruch Proportionality (Feng26) determines structure constants governing differential operators in the arithmetic proportionality framework of Feng–Yun–Zhang (FYZ26). An initial benchmark success on an already-known family led the human authors to replace their original proofs; Aletheia then calculated general eigenweights without essential human intervention using algebraic-combinatorial techniques unfamiliar to those authors.

  • The backdrop is a relation between arithmetic volumes of Chern classes on moduli spaces of shtukas and differential operators applied to L-functions of Gross motives.
  • The later interaction card documents solutions for groups of Type A, Type C, and Type D using Atiyah–Bott localization, Schur polynomial manipulation, Frobenius character identities, and the Murnaghan–Nakayama rule.
  • The milestone is autonomous generation of the mathematical content of Feng26, not autonomous authorship or a claim of a major mathematical breakthrough.

3.2. Milestone B: Lower Bounds for Multivariate Independence Polynomials

Lower bounds for multivariate independence polynomials and their generalisations (LeeSeo26) develops inequalities for independent sets and extensions modeling multiple interacting particle types. After Deep Think proved an initial inequality generalizing Sah–Sawhney–Stoner–Zhao, Aletheia supplied a high-level strategy for the harder extension, which the human authors converted into a rigorous proof.

  • Independent sets model configurations in which neighboring sites cannot both be occupied. The extension permits two types whose members do not repel members of the other type.
  • Aletheia’s contributions included dual sets, their log convexity, reduction techniques, and key lemmas for semiproper colourings.
  • The collaboration reversed the usual decomposition workflow: AI supplied the broad strategy, while humans supplied rigorous execution and sometimes independently reproved suggested statements.
  • Other models could solve the initial query, but Aletheia was the only tested model whose output on the harder query the authors found useful.

3.3. Milestone C: The Erdős Problems

From December 2–9, 2025, Aletheia with a November 2025 Gemini 3 base model attempted all 700 Erdős problems then marked Open at https://www.erdosproblems.com/. Its verifier retained 212 potentially correct responses; human screening and domain-expert review, including occasional correction of minor inaccuracies, identified 63 technically correct solutions, only 13 of which addressed the intended questions, alongside 12 ambiguous candidates.

  • Autonomous Resolution: Erdős-652 and Erdős-1051 received complete, mathematically substantive solutions considered novel after literature review.
  • Partial AI Solution: Erdős-654 and Erdős-1040 received novel solutions to parts of multi-question problems. These join the two complete resolutions in the paper’s count of four autonomous solutions.
  • Independent Rediscovery: Erdős-397, Erdős-659, Erdős-935, and Erdős-1089 had existing literature solutions found afterward. Reasoning-log inspection did not show direct retrieval, but cannot exclude indirect exposure during pretraining.
  • Literature Identification: Erdős-333, Erdős-591, Erdős-705, Erdős-992, and Erdős-1105 were already explicitly resolved in the literature despite their database status.
  • The novelty classification is an upper bound subject to further literature discoveries. None of the four novel resolutions individually met the authors’ research-paper threshold, although a collaborative generalization of Erdős-1051 produced BKKKZ26.

The 13 meaningful Erdős responses comprise two complete autonomous resolutions, two partial novel solutions, four rediscoveries, and five literature identifications. Only the first two categories contribute to the count of four novel autonomous resolutions; asterisks identify cases independently obtained elsewhere after the initial evaluations but before publication.

Taxonomy and problem identifiers for Aletheia's meaningful Erdős results.
Taxonomy and problem identifiers for Aletheia's meaningful Erdős results.

3.4. Bounds for Polynomial Dyadics

For Strongly Polynomial Policy Iteration for L∞ Robust MDPs (ACGKMP26), Aletheia improved a number-theoretic ingredient bounding how many dyadic intervals contain specified bounded combinations of numbers. Human arguments and an IMO Gold Deep Think solution already sufficed for the algorithmic application, but Aletheia’s independent use of Siegel’s Lemma yielded the strongest bound among the reported attempts and was adopted in the paper.

  • The contribution improved an intermediate proposition rather than autonomously generating the overall Robust Markov Decision Processes research project.
  • The present paper does not give the bound’s explicit formula.

4. FirstProof

FirstProof provided ten research-level questions curated by academic mathematicians independently of AI companies, largely technical lemmas drawn from their active research. The problems were released on February 5, 2026, with official solutions published at the deadline of 11:59pm PST on February 13, 2026, creating a limited evaluation window before solution contamination.

  • All questions had already been solved by mathematicians; their solutions were not online, except that a sketch for Problem 1 could be found.
  • The benchmark addresses the single-use nature of internet-equipped research evaluation: once a solution is public, later systems may retrieve it rather than independently solve the problem.

4.1. Aletheia’s results on FirstProof

Aletheia ran twice on FirstProof using January 2026 and February 2026 Gemini base models. Both runs received unmodified problem statements copied from the FirstProof LaTeX file, and a predetermined verification-and-extraction prompt filtered outputs and produced LaTeX without intermediate manual editing; Table 3 reports expert judgments for the best of the two candidates.

  • P2, P5, P7, P9, and P10 received unanimous favorable assessments, while P8 received 5/7. P1, P3, P4, and P6 produced no solution in either run.
  • P7 had previously been advertised as open, but Cappell–Weinberger–Yan had solved it before the benchmark release. The authors regard Aletheia’s distinct solution as worthy of separate documentation in FK26, not a first resolution.
  • The six-problem result uses the standard of being publishable after minor revisions and does not mean six error-free solutions from a single run.

Best-of-2 outputs received unanimous favorable judgments on P2, P5, P7, P9, and P10, with vote totals 4/4, 4/4, 3/3, 4/4, and 2/2 respectively. P8 received 5/7 and remains qualified as Correct?; four other problems had no output.

Best-of-2 FirstProof outcomes and expert correctness assessments.
Best-of-2 FirstProof outcomes and expert correctness assessments.

4.2. Comparisons

The paper uses out-of-the-box solutions to P9 and P10 by GPT 5.2 Pro, and to P10 by Gemini 3 Deep Think, as public-model baselines. It reports that OpenAI claimed six highly likely internal-model solutions, but P2 was flawed and the work involved undisclosed human guidance; it also notes an autonomously generated P6 solution from Cursor researchers.

  • These are not controlled head-to-head comparisons: model access, human intervention, and computation were not standardized.
  • The account of other submissions is qualified as the authors’ understanding of expert assessments, rather than an independent comprehensive audit.

4.3. Inference computation

Figure 5 normalizes inference costs for the six FirstProof candidate-producing problems by the cost of the earlier Erdős-1051 solution. Every candidate cost more than that reference, and P7 exceeded previously observed scales by an order of magnitude; because the base models differed, the comparison provides only a rough indication of difficulty.

  • The plotted costs cover P2, P5, P7, P8, P9, and P10 for both runs, not the four problems that produced no solution.
  • Dynamic inference expenditure is an agent-specific proxy, not a model-independent measure of mathematical difficulty.

Both runs spend more than the Erdős-1051 reference cost on every plotted FirstProof problem, with P7 the most expensive by a clear margin. The graph shows heterogeneous dynamic compute demands, but different base models prevent a controlled cross-benchmark comparison.

FirstProof inference costs normalized by the earlier Erdős-1051 solution cost.
FirstProof inference costs normalized by the earlier Erdős-1051 solution cost.

5. Analysis and Discussion

The analysis compares the agentic harness with Deep Think using the same base model, then examines correctness and qualitative limitations. These comparisons favor Aletheia on selected successful tasks, but the broader Erdős deployment and individual FirstProof runs expose substantial residual error.

5.1. Ablation studies

On the 13 Erdős problems meaningfully solved by Aletheia, same-base-model Deep Think at IMO Gold scale solved eight while using roughly 2× the average compute per problem. On the research-paper prompts, total compute was comparable, but Deep Think failed all three Feng26 queries and the crucial second LeeSeo26 query.

  • Table 4 records failures on Erdős-591, Erdős-705, Erdős-992, Erdős-1040, and Erdős-1105.
  • Table 5 records success for FYZ26, essential success for BKKKZ26 despite imperfect interaction replication, and a satisfactory but weaker bound for ACGKMP26.
  • The evaluation set was selected from Aletheia’s successes. It supports a harness advantage on these tasks, not an unbiased estimate of improvement across research problems or an isolated causal effect of verification.

Deep Think succeeds on eight of the 13 Erdős problems selected from Aletheia's successes and fails on 591, 705, 992, 1040, and 1105. Its roughly doubled average compute provides evidence that Aletheia's advantage on these selected tasks is not simply the result of a larger inference budget.

Same-base-model Deep Think results on the 13 Erdős problems meaningfully solved by Aletheia.
Same-base-model Deep Think results on the 13 Erdős problems meaningfully solved by Aletheia.

The binary summary records Deep Think success for FYZ26, BKKKZ26, and ACGKMP26, and failure for Feng26 and LeeSeo26. The body supplies necessary qualifications: BKKKZ26 was not exactly replicated, ACGKMP26's bound was weaker, and LeeSeo26's initial query succeeded while its crucial extension failed.

Deep Think's reproduction of the AI assistance underlying five research papers.
Deep Think's reproduction of the AI assistance underlying five research papers.

5.2. Accuracy

Among 200 definitively adjudicated Erdős candidates, 137 were fundamentally flawed (68.5%), 63 were technically correct (31.5%), and only 13 were meaningfully correct (6.5%). On FirstProof, both runs answered the same six questions, but the best-of-2 result combined runs containing a misinterpretation of P5, a critically flawed P7 solution, and an inadequate P8 solution.

  • The Erdős percentages are conditional on the adjudicated candidate pool, not percentages of all 700 attempted problems; 12 ambiguous responses were excluded.
  • Table 7 shows P2, P9, and P10 correct in both runs. Run A solved P5 but failed substantively on P7 and P8; Run B misinterpreted P5 but solved P7 and received a qualified favorable assessment on P8.
  • The authors regard P7 as the only publication-grade FirstProof result, giving 1/10 for that criterion. Most benchmark problems were intended as technical lemmas rather than standalone papers.
  • The Erdős evaluation used the weaker November 2025 model, so its accuracy cannot be directly attributed to the later models used for FirstProof.

Of 200 adjudicated candidate responses, 63 are technically correct but only 13 meaningfully answer the intended question. Meaningful correctness is a subset, not an additional category; the denominator excludes ambiguous candidates and problems without retained responses.

Human-graded accuracy of the 200 definitively adjudicated Erdős candidates.
Human-graded accuracy of the 200 definitively adjudicated Erdős candidates.

Run-level results show why best-of-2 reporting needs qualification: P5 is misinterpreted in Run B, P7 critically flawed in Run A, and P8 inadequate in Run A. P2, P9, and P10 are correct in both runs, while the same four problems produce no output in either.

Individual FirstProof outcomes for Aletheia Run A and Run B.
Individual FirstProof outcomes for Aletheia Run A and Run B.

5.3. Weaknesses of AI

The authors describe autonomous results as relatively short and elementary compared with typical human research papers, often relying on technical manipulation or broad knowledge retrieval rather than clearly demonstrated mathematical creativity. Aletheia remains error-prone despite verification, favors easy but unintended interpretations of ambiguous questions, and can fabricate or misrepresent results from legitimate references.

  • The 50 technically correct but vacuous Erdős answers illustrate the interpretation problem, which the paper relates to specification gaming and reward hacking.
  • The distinction between retrieval breadth and creativity is qualitative and explicitly subjective, not a measured causal finding.

6. Representing AI contributions to mathematics

Evaluating frontier mathematics requires scarce domain expertise, and even a formally correct proof does not establish novelty or significance. The paper argues that this evaluation gap, combined with publicity incentives, makes transparent reporting of AI contributions necessary and motivates a community-led taxonomy rather than a company-defined universal standard.

6.1. Autonomous Mathematics Levels

The proposed taxonomy has two independent axes: autonomy levels H, C, and A, and significance levels 0–4. Table 8 distinguishes primarily human work, essential human-AI collaboration, and essentially autonomous mathematical content; Table 9 distinguishes negligible novelty, minor novelty, Publication Grade, Major Advance, and Landmark Breakthrough.

  • 6.1.1. Discussion and Examples for Level of Autonomy. Feng26 is A; LeeSeo26 and BKKKZ26 are C; FYZ26 and ACGKMP26 are H. AI-generated inspiration alone remains H when the AI does not itself supply substantive mathematical content.
  • 6.1.2. Discussion and Examples for Level of Significance. The text classifies Erdős-652, 654, 935, and 1040 as A0; Erdős-1051 as A1; Feng26 as A2; LeeSeo26 and BKKKZ26 as C2; and FYZ26 and ACGKMP26 as H2.
  • Level 2 deliberately spans a wide range of publishable mathematics. A2 or C2 therefore does not establish parity with human mathematicians, and no reported result reaches Level 3 or Level 4.
  • Table 1 lists the Level 2 works as submitted for publication, not as having completed peer review. Mathematical significance must be judged by experts in the relevant domain.

The autonomy axis distinguishes auxiliary AI contributions, mutually essential human-AI contributions, and AI generation of the core mathematical content without essential human intervention. Even the autonomous category retains human responsibility for the final paper, separating content generation from accountability.

Autonomy levels H, C, and A for human and AI mathematical contributions.
Autonomy levels H, C, and A for human and AI mathematical contributions.

6.2. Documenting Human-AI Collaboration

Human-AI Interaction Cards document who supplied questions, ideas, proofs, and refinements. The cards for Feng26, LeeSeo26, and BKKKZ26 distinguish autonomous solutions from AI-generated outlines and non-rigorous generalization heuristics, with raw interactions linked at https://github.com/google-deepmind/superhuman/tree/main/aletheia.

  • For Feng26, humans posed successive Type A, Type C, and Type D eigenweight queries; Aletheia supplied the mathematical solutions.
  • For LeeSeo26, Aletheia supplied the crucial outline for semiproper colourings, while humans developed details and simpler proofs.
  • For BKKKZ26, Aletheia solved Erdős-1051, Deep Think suggested a weighted-product generalization, and humans weakened the hypothesis from lim = ∞ to lim sup = ∞, completed the proof, and supplied optimality examples.
  • The proposed transparency baseline for C or A results is disclosure of at least the essential raw prompts and outputs containing new AI-generated insights. The exact standards are left to the mathematical community.

7. Reflections on the Impact of AI in Mathematics

The authors interpret the results through comparative advantages rather than replacement of mathematicians: models have shallower specialist knowledge but unusually broad recall and are not constrained by human time and attention in the same way. The eigenweight case illustrates cross-field technique transfer, while the Erdős study illustrates how overlooked questions can be tractable without representing major advances.

  • These reflections do not establish that AI has matched human mathematical capability or guarantee that it will do so.

The related-work discussion covers advances in mathematical reasoning, concrete AI-assisted research results, and different human-AI collaboration architectures. It connects these developments to the need for transparent documentation, evaluation, and communication of AI-generated mathematics.

Advances in Mathematical Reasoning

The paper connects recent IMO gold-medal performance and IMO-Proof Bench results with newer research-oriented evaluations such as FrontierMath and unsolved-question benchmarks. These evaluations extend the focus beyond competition performance to advanced mathematical reasoning and research questions.

AI-Assisted Research Results

Examples of earlier AI-assisted mathematics include extremal descendant integrals on moduli spaces, motivic classes of genus zero maps, and point convergence of Nesterov’s accelerated gradient method. The cited literature also covers discrete analysis, convex analysis, combinatorics, accessible conjectures, and exploration across mathematics-adjacent disciplines.

Frameworks for Human-AI Collaboration

The paper distinguishes self-verifying and refining agent architectures, exploration systems such as AlphaEvolve, and interactive protocols for integrating AI into research practice. It describes adaptation to mathematical research as an early-stage process and calls for transparent reporting of AI contributions.

9. Conclusion

The conclusion presents specialized natural-language reasoning agents as tools for mathematicians. Natural-language systems still require correction of errors and hallucinations, while formal verification systems cannot yet readily formulate many frontier research questions; Aletheia’s informal verification is therefore an aid to research, not a guarantee of correctness or a substitute for expert responsibility.

Appendix

  • A. Case Study: FutureMath Basic presents FM-Grad-011 for q ≥ 3, N ≥ 2, and β > 0, asking when the mean-field Potts partition function has limiting free energy β/(2q) + log q. Under Aletheia solution: and the title Large Deviation Analysis of the Phase Transition in the Mean-Field Potts Model, the model uses empirical measures, Sanov’s Theorem, and Varadhan’s Lemma, reduces competing maxima to a one-dimensional optimization, and reports (\beta_{\max}=\frac{2(q-1)}{q-2}\log(q-1)). Introduction and Formulation via Empirical Measures establishes the variational expression; The Uniform Distribution and Symmetry Breaking identifies competing distributions; Determination of the Critical Inverse Temperature uses stationarity and equal-height conditions; and Conclusion states that the uniform distribution remains a global maximizer through the critical value but not above it.
  • B. Case Study: FutureMath Basic repeats FM-Grad-011 and the same titled Aletheia solution, explicitly identifying it as an example solved with the search tool. Its empirical-measure formulation, comparison of uniform and symmetry-broken distributions, and stationarity and equal-height calculation give the same critical inverse temperature and the same conclusion about the equality below and above that threshold. This is a repeated case study in the source, not a second distinct benchmark success.
  • C. Case Study: Solving IMO 2025 Problem 6 concerns a 2025 × 2025 grid with one uncovered square per row and column and a reported minimum of 2112 rectangular tiles. January 2026 Deep Think at scale 2¹², without internet or external tools, supplied a Theoretical Lower Bound based on increasing and decreasing subsequences and an Optimal Tiling Construction with 1936 central tiles and 176 edge tiles, but received an unofficial estimate of only 1–3 points because it relied on unproved, likely nonexistent results. A human-supplied refinement prompt at scale 2⁸ elicited the solution the authors report as correct: The Rectilinear Partition Formula counts corners, Segment Bounds and Chords bounds internal segments, Forcing Penalties via Monotonic Subsequences supplies the lower bound, and The Optimal Construction matches it with the final answer 2112; the reproduced output contains typographical inconsistencies.
  • D. IMO 2024 Case Studies reports PB-Advanced-021, a robustified Problem 3 variant, at scale 2⁷ with a minor mistake: the model’s justification of a uniform frequency bound required a finite-set argument supplemented by its earlier bound for large values. The Model-Generated Solution reformulates the generation rule, bounds frequencies of large numbers, characterizes infinitely recurring elements, establishes strict alternation, and uses bounded count differences and finite-state deterministic transitions to prove eventual periodicity. PB-Advanced-023, a robustified Problem 5 variant on a 3002 × 3001 table with 3000 traps, was solved at scale 2⁸: its lower-bound adversary forces at least two penalties, while Safe Probing and the “Drop to Finish” maneuver guarantee at most two, giving the final answer n = 3; On Solution Novelty contrasts this state-tracking approach with published geometric strategies, but possible training-data contamination limits claims of independent discovery.

Brief Thoughts

The same-base-model comparisons provide evidence that the complete harness improves selected research outcomes without relying solely on a larger compute budget, but they do not isolate which component causes the improvement. The gap between technical correctness and intended-problem correctness in the Erdős study, together with FirstProof’s run-level failures, remains the clearest limitation. Aletheia is best understood as a selective research-assistance system with explicit reporting of autonomy, significance, and failure—not a generally reliable autonomous mathematician.