arXiv:2512.18256cs.AIcs.LO2025-12被引 3

构建数学领域全覆盖的定理证明基准,揭示大模型在跨领域推理上的严重短板。

MSC-180: A Benchmark for Automated Formal Theorem Proving from Mathematical Subject Classification

  • 基于MSC2020分类体系,设计180个跨学科形式化证明题
  • 顶尖模型在32次尝试下仅18.89%通过率,研究生级题目通过率显著更低
  • 引入变异系数量化性能波动,暴露模型依赖训练数据模式匹配

自动化定理证明(ATP)是人工智能实现形式化推理与验证的核心方向,在推动机器智能发展方面具有重要意义。然而,当前基于大语言模型(LLM)的定理证明系统存在领域覆盖有限、数学推理泛化能力弱等局限。为此,我们提出MSC-180,一个基于MSC2020数学分类体系的评估基准。该基准包含180个形式化验证问题,每60个数学分支各选取3个高级问题,涵盖本科至研究生水平。每个问题均经领域专家多轮验证与修正,确保形式准确性。在pass@32设置下对主流LLM定理证明器进行评估,结果显示最优模型整体通过率仅为18.89%,存在明显领域偏差(最高领域覆盖41.7%)和难度差距(研究生级问题通过率显著偏低)。为进一步量化不同数学领域间的性能差异,我们引入变异系数(CV)作为评估指标,观测到的CV值为统计高变异性阈值的4-6倍,表明当前模型仍依赖训练语料中的模式匹配,缺乏可迁移的推理机制与系统泛化能力。MSC-180及其多维评估框架为推动具备真实数学推理能力的下一代AI系统发展提供了区分性强且系统化的基准。

原文摘要 · Abstract (English)

Automated Theorem Proving (ATP) represents a core research direction in artificial intelligence for achieving formal reasoning and verification, playing a significant role in advancing machine intelligence. However, current large language model (LLM)-based theorem provers suffer from limitations such as restricted domain coverage and weak generalization in mathematical reasoning. To address these issues, we propose MSC-180, a benchmark for evaluation based on the MSC2020 mathematical subject classification. It comprises 180 formal verification problems, 3 advanced problems from each of 60 mathematical branches, spanning from undergraduate to graduate levels. Each problem has undergone multiple rounds of verification and refinement by domain experts to ensure formal accuracy. Evaluations of state-of-the-art LLM-based theorem provers under the pass@32 setting reveal that the best model achieves only an 18.89% overall pass rate, with prominent issues including significant domain bias (maximum domain coverage 41.7%) and a difficulty gap (significantly lower pass rates on graduate-level problems). To further quantify performance variability across mathematical domains, we introduce the coefficient of variation (CV) as an evaluation metric. The observed CV values are 4-6 times higher than the statistical high-variability threshold, indicating that the models still rely on pattern matching from training corpora rather than possessing transferable reasoning mechanisms and systematic generalization capabilities. MSC-180, together with its multi-dimensional evaluation framework, provides a discriminative and systematic benchmark for driving the development of next-generation AI systems with genuine mathematical reasoning abilities.

定理证明数学推理基准测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。