arXiv:2605.00674cs.CL2026-05被引 71

构建持续更新的数学推理评估平台,突破静态基准局限。

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

论文配图:Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
图 1 · 摘自论文原文
  • 打造动态数学评估平台,覆盖竞赛、科研与形式化证明任务
  • GPT-5.5在2026美奥赛达98%准确率,研究题74%正确率
  • 适合追踪大模型数学能力演进的研究者与开发者

大型语言模型(LLMs)正日益成为强大的数学合作者,但静态基准已无法有效评估其进展:这些基准范围狭窄、迅速饱和且极少更新,难以可靠比较模型性能或跟踪长期进步。为此,我们需要持续维护的评估平台——能够跨多个基准运行、聚合并分析结果,以全面呈现模型在广阔领域中的表现。本文在原始MathArena基准基础上,将其范围从最终答案的奥数题扩展为持续更新的数学推理评估平台。如今,MathArena涵盖更广泛的任务,包括基于证明的竞赛、arXiv上的研究级问题以及Lean中的形式化证明生成。我们为所有模型制定清晰的评估协议,并随模型能力提升定期设计新基准,确保平台持续具有挑战性。值得注意的是,最强模型GPT-5.5在2026年美国数学奥林匹克竞赛中达到98%准确率,在研究级问题上达到74%准确率,表明前沿模型已能轻松应对极富挑战性的数学问题。这凸显了像MathArena这样的持续维护评估平台对于追踪LLMs在数学推理领域快速进展的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often narrow in scope, quickly saturated, and rarely updated. This makes it hard to compare models reliably and track progress over time. Instead, we need evaluation platforms: continuously maintained systems that run, aggregate, and analyze evaluations across many benchmarks to give a comprehensive picture of model performance within a broad domain. In this work, we build on the original MathArena benchmark by substantially broadening its scope from final-answer olympiad problems to a continuously maintained evaluation platform for mathematical reasoning with LLMs. MathArena now covers a much wider range of tasks, including proof-based competitions, research-level arXiv problems, and formal proof generation in Lean. Additionally, we maintain a clear evaluation protocol for all models and regularly design new benchmarks as model capabilities improve to ensure that MathArena remains challenging. Notably, the strongest model, GPT-5.5, now reaches 98% on the 2026 USA Math Olympiad and 74% on research-level questions, showing that frontier models can now comfortably solve extremely challenging mathematical problems. This highlights the importance of continuously maintained evaluation platforms like MathArena to track the rapid progress of LLMs in mathematical reasoning.

数学推理评估平台大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。