用AI代理团队协助数学研究,自动完成从探索到证明的全过程。
MechMath Agent Team: LLM Driven Agents for Mathematical Research

- 分三层架构分离控制、执行与增强任务,兼顾逻辑严谨与研究灵活性。
- 两个月内解决11个数论等领域的开放问题,产出形式化认证的证明。
- 适合数学研究者、AI推理系统开发者参考,推动智能科研协作。
人工智能推理已成为当代人工智能的核心焦点,主要得益于大语言模型的成功。然而,数学研究具有非线性推导路径、严格逻辑要求和长期探索周期等特点,对现有推理系统构成严峻挑战。为此,我们提出机械数学代理团队(MechMath Agent Team, MMAT),一种由大语言模型驱动的代理系统,可作为数学研究全周期的协作者。设计三重分层架构,将系统职责解耦为控制、执行与增强三个层面,实现严格逻辑控制与开放式研究敏捷性的平衡。在此框架下,构建三个专用代理:知识库管理器、自然语言证明器与形式语言证明器,形成闭环协作,生成形式化认证的数学证明。在数论、代数复杂度理论、微分代数、算子代数及不等式等领域的开放问题上进行评估。经过两个月部署,成功解决11个问题,验证了其在完整研究周期中作为协作者的能力。贡献包括:通用的多代理数学推理解耦架构、MMAT系统的具体实现,以及在多样化开放问题上的实证验证。
原文摘要 · Abstract (English)
AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathematical research, which is characterized by non-linear derivation paths, rigorous logical requirements, and protracted exploration cycles, poses severe challenges for existing reasoning systems. To overcome these limitations, we present the MechMath Agent Team (MMAT), which is a large language model driven agent designed to serve as a co-pilot throughout the full cycle of mathematical research. We design a tripartite Harness Architecture that decouples system responsibilities into Control, Execution, and Augmentation planes, thereby reconciling rigorous logical control with the agility demanded by open-ended research. Building upon this framework, we instantiate three specialized agents: a Knowledge Base Manager, a Natural Language Prover, and a Formal Language Prover, all operating in a closed loop to produce formally certified mathematical proofs. We evaluate MMAT on open problems in Number Theory, Algebraic Complexity Theory, Differential Algebra, Operator Algebra, and Inequalities. Across a two-month deployment, 11 problems have been solved, demonstrating its capacity to act as a co-pilot throughout the entire research cycle. The contributions are threefold: a general decoupled Harness Architecture for multi-agent mathematical reasoning, its concrete instantiation in the MMAT system, and empirical validation on a diverse suite of open problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。