arXiv:2504.01995cs.AIcs.LG2025-04被引 26

评测大模型解奥数题真本事,发现多数正确答案靠套路而非推理。

Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics

  • 设计自动评估框架,从逻辑严谨性角度分析模型推理过程。
  • 仅12%的解题过程符合数学证明标准,多数正确答案实为巧合。
  • 适合关注AI数学能力真实水平的研究者与教育科技开发者。

大型语言模型(LLMs)在数学推理任务中表现出显著进展。然而,当前评估基准多关注最终答案的准确率,忽视了解题过程中的逻辑严谨性。本文通过定性和定量的人类评估,研究了LLMs生成的数学证明,并开发了一种自动评估其推理能力的框架。研究发现,现有LLMs在解决高难度奥数题方面存在明显不足,常无法区分正确的数学推理与明显错误的推导。分析表明,少数看似正确的答案往往源于模式识别或启发式捷径,而非真正的数学推理。这些结果揭示了大模型在高级数学推理上与人类专家之间的巨大差距,强调应建立以推理过程合理性为核心的新评估标准,而非仅关注最终答案的正确性。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have shown impressive progress in mathematical reasoning tasks. However, current evaluation benchmarks predominantly focus on the accuracy of final answers, often overlooking the crucial logical rigor for mathematical problem solving. The claim that state-of-the-art LLMs can solve Math Olympiad-level problems requires closer examination. To explore this, we conducted both qualitative and quantitative human evaluations of proofs generated by LLMs, and developed a schema for automatically assessing their reasoning capabilities. Our study reveals that current LLMs fall significantly short of solving challenging Olympiad-level problems and frequently fail to distinguish correct mathematical reasoning from clearly flawed solutions. Our analyses demonstrate that the occasional correct final answers provided by LLMs often result from pattern recognition or heuristic shortcuts rather than genuine mathematical reasoning. These findings underscore the substantial gap between LLM performance and human expertise in advanced mathematical reasoning and highlight the importance of developing benchmarks that prioritize the soundness of the reasoning used to arrive at an answer rather than the mere correctness of the final answers.

数学推理大模型评估奥数题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。