用实时更新的数学竞赛题评估大模型,避免答案记忆,检验推理与证明能力。
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
- 基于新发布的数学竞赛题构建实时评测框架,杜绝数据泄露。
- 顶级模型在IMO 2025中证明题得分接近40%,展现推理潜力。
- 首个评估模型写证明能力的基准,适合关注数学推理的研究者。
大语言模型在数学推理上的快速进步推动了其在数学基准上的显著提升。然而,许多常用评测数据集(如AIME 2024)广泛存在于网络,难以区分真实推理与潜在记忆。此外,这些基准未评估证明写作能力,而后者对多数数学任务至关重要。为此,我们提出MathArena,核心思路是:定期举办的数学竞赛可提供持续高质量、高难度题目,用于大模型的实时评估。通过在新题发布后立即评测,有效消除数据污染风险。我们发现AIME 2024存在明显污染迹象。但在更难的竞赛如CMIMC 2025上,顶尖模型展现出强大推理能力。MathArena也是首个专门评估证明写作能力的基准。在IMO 2025中,顶级模型得分略低于40%,表明虽有进展但仍存巨大提升空间。截至目前,已对50余种模型在7场竞赛中进行评测,共162道题。作为动态基准,MathArena将持续追踪模型在新竞赛中的表现,确保数学推理评估的严谨性与前沿性。
原文摘要 · Abstract (English)
The rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used evaluation datasets (e.g., AIME 2024) are widely available online, making it difficult to disentangle genuine reasoning from potential memorization. Furthermore, these benchmarks do not evaluate proof-writing capabilities, which are crucial for many mathematical tasks. To address this, we introduce MathArena, a new benchmark based on the following key insight: recurring math competitions provide a stream of high-quality, challenging problems that can be used for real-time evaluation of LLMs. By evaluating models as soon as new problems are released, we effectively eliminate the risk of contamination. Using this framework, we find strong signs of contamination in AIME 2024. Nonetheless, evaluations on harder competitions, such as CMIMC 2025, demonstrate impressive reasoning capabilities in top-performing models. MathArena is also the first benchmark for proof-writing capabilities. On IMO 2025, top models achieve slightly less than 40%, demonstrating both notable progress and significant room for improvement. So far, we have evaluated over $50$ models across seven competitions, totaling $162$ problems. As an evolving benchmark, MathArena will continue to track the progress of LLMs on newly released competitions, ensuring rigorous and up-to-date evaluation of mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。