评测大模型在高阶数学证明生成与验证中的表现,发现仍存在巨大提升空间。
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

- 构建覆盖本科至博士资格考级别的证明题库,支持细粒度推理评估。
- 最优模型在证明生成任务上准确率仅达75.8%(UGD)和66.1%(QE)。
- 模型验证能力弱,误判率高,尤其难以识别关键错误,适合研究者参考。
大型语言模型(LLMs)在中学及奥数级别数学中表现优异,但在高阶数学领域的能力尚不明确。现有基准在覆盖面和评估粒度上均不足:学科覆盖有限,且多依赖最终答案正确性或粗略判断,未能充分评估推理过程的有效性。为此,我们提出AdvancedMathBench,一个专为评估高阶数学推理能力设计的基准套件。其核心证明生成基准ProverBench包含296道题目,涵盖本科及博士资格考试难度。为实现可靠评估,我们开发了基于大规模专家标注训练的自动验证流水线,可输出正确性判定与细粒度错误分析,与人类专家在保留证明轨迹上的判断高度一致。此外,我们引入VerifierBench,包含888条模型生成的证明轨迹及其专家真实标签,用于评估模型判断证明有效性及提供合理验证理由的能力。实验表明,AdvancedMathBench对前沿模型仍具挑战:在证明生成任务中,表现最佳的GPT-5.5-xhigh模型在UGD和QE划分上的得分分别为75.8和66.1;在验证任务中,最优模型仅获得65.1的平衡F1,且普遍真负率偏低,表明关键错误检测仍是主要瓶颈。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-generation benchmark, ProverBench, contains 296 problems spanning undergraduate and doctoral qualifying-exam levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 75.8 and 66.1 on the UGD and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model attains a Balanced F1 of only 65.1, and models generally exhibit low true negative rates, suggesting that critical error detection remains a major bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。