arXiv:2602.10177cs.LGcs.AI2026-02被引 23

AI自主生成数学研究论文并解决多个开放问题。

Towards Autonomous Mathematics Research

  • 构建端到端的数学研究代理Aletheia,通过迭代生成与验证推进证明。
  • 自主解决4个开放问题,产出无须人工干预的研究论文。
  • 适合关注AI辅助科研、数学推理与人机协作的学者与开发者。

基础模型的进展已使推理系统达到国际数学奥林匹克竞赛金牌水平。然而,从竞赛级解题转向专业研究,需应对海量文献与长周期证明。本文提出数学研究代理Aletheia,通过自然语言实现解题的迭代生成、验证与修正。其依托增强版Gemini Deep Think处理复杂推理,引入新型推理时扩展规律以突破奥数级别限制,并深度使用工具应对数学研究复杂性。我们展示了Aletheia从奥数题到博士级习题的能力,尤为关键的是在人工智能辅助数学研究中实现三项里程碑:(a) 由AI完全自主生成的论文Feng26,计算算术几何中的特征权重结构常数;(b) 人类-人工智能协作论文LeeSeo26,证明交互粒子系统中独立集的界;(c) 对Bloom的Erdos猜想数据库中700个开放问题的半自主评估,包括4个自主求解的开放问题。为促进公众理解,我们建议量化自主性与创新性标准,并提出人机互动卡片以提升透明度。结论部分反思人机协同在数学中的前景,项目所有提示与模型输出均开源于https://github.com/google-deepmind/superhuman/tree/main/aletheia。

原文摘要 · Abstract (English)

Recent advances in foundational models have yielded reasoning systems capable of achieving a gold-medal standard at the International Mathematical Olympiad. The transition from competition-level problem-solving to professional research, however, requires navigating vast literature and constructing long-horizon proofs. In this work, we introduce Aletheia, a math research agent that iteratively generates, verifies, and revises solutions end-to-end in natural language. Specifically, Aletheia is powered by an advanced version of Gemini Deep Think for challenging reasoning problems, a novel inference-time scaling law that extends beyond Olympiad-level problems, and intensive tool use to navigate the complexities of mathematical research. We demonstrate the capability of Aletheia from Olympiad problems to PhD-level exercises and most notably, through several distinct milestones in AI-assisted mathematics research: (a) a research paper (Feng26) generated by AI without any human intervention in calculating certain structure constants in arithmetic geometry called eigenweights; (b) a research paper (LeeSeo26) demonstrating human-AI collaboration in proving bounds on systems of interacting particles called independent sets; and (c) an extensive semi-autonomous evaluation (Feng et al., 2026a) of 700 open problems on Bloom's Erdos Conjectures database, including autonomous solutions to four open questions. In order to help the public better understand the developments pertaining to AI and mathematics, we suggest quantifying standard levels of autonomy and novelty of AI-assisted results, as well as propose a novel concept of human-AI interaction cards for transparency. We conclude with reflections on human-AI collaboration in mathematics and share all prompts as well as model outputs at https://github.com/google-deepmind/superhuman/tree/main/aletheia.

AI研究数学推理人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。