新基准挑战顶尖大模型数学推理能力,顶级模型仅答对52.4%。
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
- 构建50道原创高难度奥数题,经专家验证达国际奥赛水平
- 26个大模型测试平均分低于40%,最强模型仅52.4%正确率
- 适合研究模型推理能力提升,尤其关注计算资源与性能关系
我们提出AMO-Bench,一个高级数学推理基准,包含50道人工设计的难题,难度达到甚至超过国际数学奥林匹克(IMO)级别。现有基准多基于高中数学竞赛,但因模型表现趋于饱和(如AIME24/25),已难有效评估顶尖大语言模型(LLMs)。为解决此问题,AMO-Bench确保所有题目:(1) 经专家交叉验证,满足至少IMO难度标准;(2) 完全原创,防止数据记忆导致的性能泄露。此外,每道题仅需最终答案,支持自动且可靠的评分。在AMO-Bench上对26个LLMs的实验显示,即使最优模型准确率也仅达52.4%,多数模型低于40%。进一步分析揭示,随着测试时计算量增加,性能呈现显著提升趋势。这些结果表明当前大模型在数学推理方面仍有巨大改进空间。我们开源AMO-Bench以推动语言模型推理能力的研究。https://amo-bench.github.io/
原文摘要 · Abstract (English)
We present AMO-Bench, an Advanced Mathematical reasoning benchmark with Olympiad level or even higher difficulty, comprising 50 human-crafted problems. Existing benchmarks have widely leveraged high school math competitions for evaluating mathematical reasoning capabilities of large language models (LLMs). However, many existing math competitions are becoming less effective for assessing top-tier LLMs due to performance saturation (e.g., AIME24/25). To address this, AMO-Bench introduces more rigorous challenges by ensuring all 50 problems are (1) cross-validated by experts to meet at least the International Mathematical Olympiad (IMO) difficulty standards, and (2) entirely original problems to prevent potential performance leakages from data memorization. Moreover, each problem in AMO-Bench requires only a final answer rather than a proof, enabling automatic and robust grading for evaluation. Experimental results across 26 LLMs on AMO-Bench show that even the best-performing model achieves only 52.4% accuracy on AMO-Bench, with most LLMs scoring below 40%. Beyond these poor performances, our further analysis reveals a promising scaling trend with increasing test-time compute on AMO-Bench. These results highlight the significant room for improving the mathematical reasoning in current LLMs. We release AMO-Bench to facilitate further research into advancing the reasoning abilities of language models. https://amo-bench.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。