arXiv:2509.26076cs.CL2025-09被引 12

评测大模型在顶级数学证明上的能力,真实模拟科研环境。

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

  • 用专家设计的77道前沿数学题构建新基准
  • 模型在真实研究场景中可解决超三成难题
  • 适合关注AI数学推理与科研辅助的研究者

随着大语言模型(LLMs)数学能力提升,评估其在前沿研究级任务中的表现愈发重要。现有基准多聚焦最终答案或中学竞赛题,难以反映真实科研水平。为此,我们提出IMProofBench,一个由专家数学家设计的私有基准,包含77道同行评审的题目。每道题需完整证明,配套子问题提供最终答案,支持人工评估与自动化评分。评估环境模拟真实科研:模型在代理框架下使用网络搜索和SageMath等工具。结果表明,当前LLMs已能解决超过30%的研究级问题。IMProofBench将协同数学界持续演化,保持对下一代LLMs的评估相关性。

原文摘要 · Abstract (English)

As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of mathematical knowledge. However, existing benchmarks are limited, as they focus solely on final-answer questions or high-school competition problems. To address this gap, we introduce IMProofBench, a private benchmark consisting of 77 peer-reviewed problems developed by expert mathematicians. Each problem requires a detailed proof and is paired with subproblems that have final answers, supporting both an evaluation by human experts and a large-scale quantitative analysis through automated grading. Furthermore, unlike prior benchmarks, the evaluation setup simulates a realistic research environment: models operate in an agentic framework with tools like web search for literature review and mathematical software such as SageMath. Our results show that current LLMs can already solve a significant percentage of research-level questions. IMProofBench will continue to evolve as a dynamic benchmark in collaboration with the mathematical community, ensuring its relevance for evaluating the next generation of LLMs.

数学推理大模型评测科研辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。