让大模型从解题工具变身为能探索数学前沿的研究伙伴
From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

- 将大模型从固定问题求解转向自主探索数学前沿
- 指出现有系统在数据、结构、探索能力上的根本缺陷
- 适合关注AI辅助数学研究的学者与跨学科研究者
近年来,人工智能在数学领域(AI4Math)取得显著进展,尤其是基于大语言模型(LLM)的定理证明系统,在交互式定理证明(ITP)语言中已能高效生成形式化证明。然而,当前系统仍难以应对数学前沿研究,如发现新定理或解决开放猜想,这些问题往往开放性强、定义模糊,并涉及多层抽象。本文主张,AI4Math的下一次飞跃需实现从预设问题求解器到具备严谨形式化推理能力的研究代理的转变。本文系统梳理该领域,涵盖数据集、自动形式化与证明合成;更重要的是,识别现有系统作为数学研究代理的核心局限,从数据、关系结构、数学探索、工具生态及人机协作等多个维度展开分析,提出未来发展的战略路线图。
原文摘要 · Abstract (English)
Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures, which are often open-ended, under-specified, and involve multiple layers of abstraction. We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges with rigorous formal mathematical reasoning. In this position paper, we provide a systematic review of the field, covering datasets, auto-formalization, and proof synthesis. More importantly, we identify core limitations of existing systems in serving as mathematical research agents, examining issues across datasets, relational structure, mathematical exploration, tool ecosystem, and human-AI collaboration, outlining a strategic road-map for the future of AI4Math.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。