arXiv:2603.26996cs.AIcs.CL2026-03中稿 · ICLR被引 3

评测大模型能否写出可形式化验证的研究生级数学证明

FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?

  • 构建专用基准,要求模型输出可通过Lean 4检查器验证的正式证明
  • 最先进模型仅达33.5%正确率,且性能随难度快速下降
  • 适合关注形式化推理与大模型数学能力的研究者

我们提出FormalProofBench,一个私有基准,用于评估AI模型在研究生水平上生成可形式化验证数学证明的能力。每个任务包含自然语言问题和对应的Lean 4形式化陈述,模型需输出被Lean 4检查器接受的证明。该基准覆盖分析、代数、概率与逻辑等领域,题目来自资格考试与标准教材。我们使用代理式框架评估多种前沿模型,发现最佳基础模型准确率为33.5%,且性能随难度迅速下降。除准确率外,还提供工具使用、失败模式、成本与延迟的实证分析,全面评估前沿模型的形式定理证明能力。

原文摘要 · Abstract (English)

We present FormalProofBench, a private benchmark designed to evaluate whether AI models can produce formally verified mathematical proofs at the graduate level. Each task pairs a natural-language problem with a Lean~4 formal statement, and a model must output a Lean proof accepted by the Lean 4 checker. FormalProofBench targets advanced undergraduate and graduate mathematics, with problems drawn from qualifying exams and standard textbooks across topics including analysis, algebra, probability, and logic. We evaluate a range of frontier models with an agentic harness, and find that the best-performing foundation model achieves 33.5% accuracy, with performance dropping rapidly after that. In addition to the accuracy numbers, we also provide empirical analysis of tool-use, failure modes, cost and latency, thereby providing a thorough evaluation of the formal-theorem proving abilities of frontier models.

形式化证明数学推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。