用最新论文构建动态数学推理测试,评估大模型真实理解能力。
LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches
- 基于近期arXiv论文构建动态多选题,避免数据泄露。
- 引入13类定理分类与证明草图生成干扰项,提升评测敏感度。
- 支持抗替换验证,揭示模型是否真会推理而非记忆答案。
数学推理是人类智能的标志,大语言模型(LLMs)能否真正具备此能力仍是人工智能与认知科学的核心问题。随着LLMs被广泛集成于科研工作流,对其数学能力进行严格评估成为现实需求。现有基准受限于合成场景和数据污染。我们提出LiveMathematicianBench,一个基于训练截止后新发表arXiv论文构建的动态多选题基准,用于研究级数学推理评估。通过以最新发表的定理为依据,该基准提供了超越记忆模式的真实测试环境。基准引入13类逻辑定理分类(如蕴含、等价、存在性、唯一性),实现对不同推理形式的细粒度评估。采用基于证明草图的干扰项生成流程,利用高层次证明策略构造看似合理但无效的答案选项,增强对真实理解力的敏感性。同时引入抗替换机制,区分答案识别与实质性推理。评估显示该基准尚未饱和:表现最佳的Gemini-3.1-pro-preview仅达43.5%。在抗替换评估下,准确率急剧下降:GPT-5.4最高为30.6%,而Gemini-3.1-pro-preview降至17.6%,低于20%随机基线。双模式协议表明,访问证明草图能持续提升准确率,说明模型可利用高层证明策略进行推理。总体而言,LiveMathematicianBench提供了一个可扩展、抗污染的研究级数学推理评估平台。
原文摘要 · Abstract (English)
Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. As LLMs are increasingly integrated into scientific workflows, rigorous evaluation of their mathematical capabilities becomes a practical necessity. Existing benchmarks are limited by synthetic settings and data contamination. We present LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs. By grounding evaluation in newly published theorems, it provides a realistic testbed beyond memorized patterns. The benchmark introduces a thirteen-category logical taxonomy of theorem types (e.g., implication, equivalence, existence, uniqueness), enabling fine-grained evaluation across reasoning forms. It employs a proof-sketch-guided distractor pipeline that uses high-level proof strategies to construct plausible but invalid answer choices reflecting misleading proof directions, increasing sensitivity to genuine understanding over surface-level matching. We also introduce a substitution-resistant mechanism to distinguish answer recognition from substantive reasoning. Evaluation shows the benchmark is far from saturated: Gemini-3.1-pro-preview, the best model, achieves only 43.5%. Under substitution-resistant evaluation, accuracy drops sharply: GPT-5.4 scores highest at 30.6%, while Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline. A dual-mode protocol reveals that proof-sketch access yields consistent accuracy gains, suggesting models can leverage high-level proof strategies for reasoning. Overall, LiveMathematicianBench offers a scalable, contamination-resistant testbed for studying research-level mathematical reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。