构建IMO级别数学推理评测集,推动大模型突破奥数难题
Towards Robust Mathematical Reasoning
- 设计双阶评测体系:短答案与证明题并重
- Gemini模型在奥数题上达80%准确率,证明题超非Gemini模型42.4%
- 开发自动评分系统,支持长文本答案评估
为提升基础模型的数学推理能力,现有评估或过于简单,或仅关注简短答案。为此,我们提出IMO-Bench,一套由顶级专家评审的高级推理评测集,专攻国际数学奥林匹克(IMO)级别题目。IMO-AnswerBench包含400道多样化的奥数题,验证短答案;IMO-Proof Bench则评估证明写作能力,涵盖基础与高阶奥数题,并提供详细评分标准以实现自动评分。该评测在Gemini Deep Think(Luong和Lockhart,2025)中助力实现2025年IMO金牌级表现,模型在IMO-AnswerBench上达到80.0%,在高级证明题上达65.7%,分别领先最佳非Gemini模型6.9%与42.4%。我们还验证了基于Gemini推理的自动评分器与人工评分高度一致,并构建了含1000份人工评分的IMO-GradingBench,以推动长文本答案的自动化评估。我们希望IMO-Bench能促进社区在鲁棒数学推理上的进展,并已公开发布于https://imobench.github.io/。
原文摘要 · Abstract (English)
Finding the right north-star metrics is highly critical for advancing the mathematical reasoning capabilities of foundation models, especially given that existing evaluations are either too easy or only focus on getting correct short answers. To address these issues, we present IMO-Bench, a suite of advanced reasoning benchmarks, vetted by a panel of top specialists and that specifically targets the level of the International Mathematical Olympiad (IMO), the most prestigious venue for young mathematicians. IMO-AnswerBench first tests models on 400 diverse Olympiad problems with verifiable short answers. IMO-Proof Bench is the next-level evaluation for proof-writing capabilities, which includes both basic and advanced IMO level problems as well as detailed grading guidelines to facilitate automatic grading. These benchmarks played a crucial role in our historic achievement of the gold-level performance at IMO 2025 with Gemini Deep Think (Luong and Lockhart, 2025). Our model achieved 80.0% on IMO-AnswerBench and 65.7% on the advanced IMO-Proof Bench, surpassing the best non-Gemini models by large margins of 6.9% and 42.4% respectively. We also showed that autograders built with Gemini reasoning correlate well with human evaluations and construct IMO-GradingBench, with 1000 human gradings on proofs, to enable further progress in automatic evaluation of long-form answers. We hope that IMO-Bench will help the community towards advancing robust mathematical reasoning and release it at https://imobench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。