用大模型自动攻克前沿数学难题,突破竞赛题局限。
AI Mathematician: Towards Fully Automated Frontier Mathematical Research
- 设计探索机制与保守验证法,应对研究难题的复杂性与严谨性要求。
- 在多个真实数学领域自主构建证明片段并发现非平凡洞察。
- 适合对数学自动推理与AI辅助科研感兴趣的学者与开发者。
大型推理模型(LRMs)近年来在数学能力上取得显著进展,但主要局限于竞赛级问题。本文提出AI Mathematician(AIM)框架,利用LRMs的推理能力支持前沿数学研究。相比竞赛题,研究问题具有内在复杂性与程序严谨性要求。为此,AIM引入探索机制以拓展解题路径,并采用悲观合理验证方法确保结果可靠性。该早期版本已展现出处理研究级任务的强大能力。我们在多个真实数学领域开展广泛实验,结果表明:AIM能自主构建大量证明内容,并在各研究方向发现非平凡洞见。这些发现凸显了LRMs在数学发现中的潜力,预示基于LRM的智能体系统未来可大幅加速数学研究进程。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have made significant progress in mathematical capabilities in recent times. However, these successes have been primarily confined to competition-level problems. In this work, we propose AI Mathematician (AIM) framework, which harnesses the reasoning strength of LRMs to support frontier mathematical research. We have identified two critical challenges of mathematical research compared to competition, {\it the intrinsic complexity of research problems} and {\it the requirement of procedural rigor}. To address these challenges, AIM incorporates two core strategies: an exploration mechanism to foster longer solution paths, and the pessimistic reasonable verification method to ensure reliability. This early version of AIM already exhibits strong capability in tackling research-level tasks. We conducted extensive experiments across several real-world mathematical topics and obtained promising results. AIM is able to autonomously construct substantial portions of proofs and uncover non-trivial insights within each research area. These findings highlight the potential of LRMs in mathematical discovery and suggest that LRM-based agent systems could significantly accelerate mathematical research in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。