arXiv:2503.21934cs.CL2025-03被引 91

测试大模型在2025年美国数学奥林匹克中的完整解题能力,发现表现远低于人类。

Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

  • 基于专家标注,评估模型对六道奥数题的完整解题过程。
  • 最强模型Gemini-2.5-Pro仅得25%,其余均低于5%。
  • 揭示训练优化导致的推理错误与伪证明问题,适合关注推理缺陷的研究者。

近期数学基准测试(如MathArena)显示,先进推理模型在AIME等竞赛中表现优异,顶级模型Gemini-2.5-Pro得分接近顶尖人类选手。然而,这些测试仅依据最终数值答案评分,忽视了严谨推理与证明生成——这正是真实数学任务的核心。为此,我们引入对复杂数学问题全解题过程的综合评估。利用专家人工标注,在2025年美国数学奥林匹克(USAMO)六道题目发布后数小时内,评估了几种前沿推理模型。结果表明,所有模型表现均显著不足:仅有Gemini-2.5-Pro获得非平凡分25%,其余模型得分均低于5%。通过对推理轨迹的详细分析,我们识别出主要失败模式,并发现训练优化策略引发多种不当伪证与逻辑漏洞。总体而言,当前大语言模型仍难以胜任严格的数学推理任务,亟需提升其形式化推理与证明生成能力。

原文摘要 · Abstract (English)

Recent math benchmarks for large language models (LLMs) such as MathArena indicate that state-of-the-art reasoning models achieve impressive performance on mathematical competitions like AIME, with the leading model, Gemini-2.5-Pro, achieving scores comparable to top human competitors. However, these benchmarks evaluate models solely based on final numerical answers, neglecting rigorous reasoning and proof generation which are essential for real-world mathematical tasks. To address this, we introduce a comprehensive evaluation of full-solution reasoning for challenging mathematical problems. Using expert human annotators, we evaluated several state-of-the-art reasoning models on the six problems from the 2025 USAMO within hours of their release. Our results reveal that all tested models struggled significantly: only Gemini-2.5-Pro achieves a non-trivial score of 25%, while all other models achieve less than 5%. Through detailed analysis of reasoning traces, we identify the most common failure modes and find several unwanted artifacts arising from the optimization strategies employed during model training. Overall, our results suggest that current LLMs are inadequate for rigorous mathematical reasoning tasks, highlighting the need for substantial improvements in reasoning and proof generation capabilities.

数学推理大模型评估证明生成USAMO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。