arXiv:2605.22763cs.AI2026-05被引 17

AI可自动证明数学难题,9个埃爾多斯问题被攻克

Advancing Mathematics Research with AI-Driven Formal Proof Search

论文配图:Advancing Mathematics Research with AI-Driven Formal Proof Search
图 1 · 摘自论文原文
  • 用大模型生成形式化证明,再由Lean系统验证
  • 自主解决9个埃爾多斯开放问题,成本仅数百美元/题
  • 适用于组合数学、图论等领域的前沿研究

大语言模型在数学推理上表现日益出色,但其不可靠性限制了在数学研究中的应用。一种缓解方案是使用大模型生成形式化证明(如Lean语言)。我们首次大规模评估该方法解决开放问题的能力。最先进代理在每题数百美元成本下,自主解决了353个埃爾多斯问题中的9个,证实了44/492个OEIS猜想,并已部署于组合数学、优化、图论、代数几何和量子光学研究中。基础代理通过交替使用大模型生成与Lean验证,复现了埃爾多斯成果,但在最难题目上成本更高。这些结果展示了人工智能辅助形式化证明的强大能力,也为有效代理设计提供了洞见。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research. A mitigation is using LLMs to generate formal proofs in languages like Lean. We perform the first large-scale evaluation of this method's ability to solve open problems. Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures, and is being deployed in combinatorics, optimization, graph theory, algebraic geometry, and quantum optics research. A basic agent alternating LLM-based generation with Lean-based verification replicated the Erdős successes but proved costlier on the hardest problems. These findings demonstrate the power of AI-aided formal proof search and shed light on the agent designs that enable it.

形式证明数学推理大模型自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。