arXiv:2607.14582cs.AI2026-07被引 1

让数学家与AI协作证明定理,实时互动推进研究。

MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research

论文配图:MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research
图 1 · 摘自论文原文
  • 人机协同构建可交互的证明蓝图,分步引导推理过程。
  • 在偏微分方程定理上,模型验证成功率低于50%。
  • 适合需要深度领域理解的数学研究者使用。

现有基于大语言模型的定理证明系统虽在形式化数学基准上表现优异,但仅能作为自主代理证明给定命题。本文提出MathCoPilot,一种人机协同的数学研究新范式:数学家主导高层次方向,AI代理在持续人机指导下完成形式化与证明细节。该系统集成三大能力:(1) 交互式工作台,支持人机通过动态证明蓝图分步协作;(2) 自适应知识库检索与集成Lean的迭代验证机制;(3) 主题驱动论文检索与自动化形式化至可验证的Lean知识库。我们使用MathCoPilot系统,对Gemini 3.1 Pro、GPT-5.4和Claude Opus 4.7等四款先进LLM在FormalMATH子集及两个需深度领域知识的真实偏微分方程定理上进行评估,考察其生成可验证Lean 4证明的能力以及识别故意错误证明的能力。结果显示,当前模型在有利自动形式化条件下可高成功率解决本科级问题,但在需真正数学理解的领域特定定理上仍面临显著挑战。

原文摘要 · Abstract (English)

Existing LLM-based theorem provers have achieved impressive results on formal mathematics benchmarks, yet they remain confined to acting as autonomous agents that prove a stated proposition. In this paper, we propose MathCoPilot, a human-in-the-loop system that embodies a new human--AI symbiotic paradigm for mathematical research, in which the mathematician steers the high-level mathematical direction while AI agents carry out the detailed formalization and proof work under continuous human guidance. MathCoPilot unifies three core capabilities: (1) an interactive workbench where the mathematician and AI agents collaborate through a living proof blueprint that decomposes a proof into navigable steps the human can directly inspect, direct, and refine; (2) automated proving skill orchestration with adaptive knowledge base search and Lean-integrated iterative verification; and (3) topic-driven paper retrieval and automated formalization into a verified Lean knowledge base. Using MathCoPilot, we systematically compare four state-of-the-art LLMs, including Gemini~3.1~Pro, GPT-5.4, and Claude~Opus~4.7, on a FormalMATH subset and on two real PDE theorems requiring deep domain expertise, evaluating their ability to produce verified Lean~4 proofs and to identify errors in deliberately incorrect proofs. Our results show that while current models can handle undergraduate-level problems with high success rates under favorable autoformalization conditions, substantial challenges remain for domain-specific theorems requiring genuine mathematical understanding.

人机协同定理证明Lean

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。