arXiv:2605.22875cs.AIcs.LG2026-05被引 3

RMA让AI能自主解决高阶数学研究难题,通过多轮协作推理与验证。

RMA: an Agentic System for Research-Level Mathematical Problems

论文配图:RMA: an Agentic System for Research-Level Mathematical Problems
图 1 · 摘自论文原文
  • 构建多角色代理系统,分步完成问题分析、文献检索与证明迭代。
  • 在10个专家级数学问题上解决8个,逻辑更严谨,可读性更强。
  • 适合需要长程推理的数学研究者,推动AI辅助科研发展。

我们提出研究数学智能体(RMA),一个用于自动化求解研究级数学问题的智能体框架。不同于聚焦竞赛数学或形式化定理证明的前期研究,RMA针对需要长周期推理、文献支撑和迭代证明优化的研究级问题。RMA将研究级证明求解分解为问题分析、文献搜索与理解、公平比较、知识库构建和证明验证等专用模块,由初始化器、提议者和验证者代理通过共享结构化记忆协同调度。在此统一框架下,各代理以多角色、多轮次方式协作生成、优化并验证候选证明,通过迭代反馈持续改进。我们在涵盖多个领域的十道专家贡献的研究级问题上评估RMA,结果表明其性能超越强基线模型(包括GPT-5.2R和Aletheia),成功解决其中八题,并生成更逻辑严密、更易读的证明。全面的消融实验进一步显示,性能提升源于结构化推理模块、迭代优化与验证器反馈的协同作用,而非单一组件。解决方案与实现代码将在论文接受后公开。

原文摘要 · Abstract (English)

We present $\textbf{Research Math Agents (RMA)}$, an agentic framework for automated reasoning on research-level mathematical problems. Unlike prior studies centered on competition mathematics or formal theorem proving, RMA targets research-level mathematical problems that require long-horizon reasoning, literature grounding, and iterative proof refinement. RMA decomposes research-level proof solving into specialized modules for problem analysis, literature search and understanding, fair comparison, knowledge-bank construction, and proof verification, all coordinated by initializer, proposer, and verifier agents through a shared structured memory. Within this unified framework, these agents operate in a multi-role, multi-round workflow, collaboratively generating, refining, and verifying candidate proofs through iterative feedback. We evaluate RMA on the First Proof benchmark, which consists of ten research-level problems contributed by expert mathematicians across diverse domains. Through comprehensive expert evaluation, RMA outperforms strong baselines on the First Proof benchmark, including GPT-5.2R and Aletheia, solving eight out of ten research problems and producing more logically sound and readable proofs. Our comprehensive ablation studies further show that performance gains arise from the interaction of structured reasoning modules, iterative refinement, and verifier-based feedback, rather than any single component. Our solutions and implementations will be made publicly available upon acceptance.

数学推理智能体系统研究辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。