RIMO为高阶数学推理设计了易评估、难解的奥数基准测试
RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning
- 将335道奥数题改写为唯一整数答案,实现确定性判错
- 456道证明题分解为子问题,自动评分验证推理过程
- 揭露顶尖大模型在真正奥数题上表现骤降,适合挑战强推理能力
随着大语言模型在GSM8K和MATH等基准上取得高分,研究界转向国际数学奥林匹克(IMO)题目以推进评估边界。然而现有奥数级基准存在实际限制,如答案格式不一需依赖模型裁判、解决方案可能有误,导致评分噪声与偏差。本文提出RIMO,一个双轨基准:RIMO-N将335道IMO题重写为仅有一个唯一整数答案,支持确定性正确性判断;RIMO-P包含456道经专家验证的证明题,拆解为子问题序列,通过自动化系统评估逐步推理过程。对十种前沿LLM(包括GPT-4o和Gemini 2.5 Flash)的测评显示,尽管这些模型在旧基准表现优异,但在RIMO上性能急剧下降,揭示当前模型与真正奥数级推理能力之间的巨大差距。RIMO提供了一个既具挑战性又易于评估的测评工具,为未来研究提供高精度标尺,明确指出了需弥补的深层推理鸿沟。
原文摘要 · Abstract (English)
As large language models (LLMs) reach high scores on established mathematical benchmarks, such as GSM8K and MATH, the research community has turned to International Mathematical Olympiad (IMO) problems to push the evaluation frontier. However, existing Olympiad-level benchmarks suffer from practical constraints that introduce grading noise and potential bias, such as heterogeneous answer formats requiring model-based judges and a reliance on potentially flawed solutions. We introduce RIMO, a two-track benchmark designed to preserve peak Olympiad difficulty while eliminating this evaluation noise. The first track, RIMO-N, rewrites 335 IMO problems to admit a single, unique integer answer, allowing for deterministic correctness checking. The second track, RIMO-P, features 456 proof problems with expert-checked solutions, which are decomposed into a sequence of sub-problems to evaluate the step-by-step reasoning process via an automated grading system. Our benchmarking of ten frontier LLMs, including GPT-4o and Gemini 2.5 Flash, reveals that while these systems excel on older benchmarks, their performance drops sharply on RIMO. These results highlight a substantial gap between current LLM capabilities and actual Olympiad-level reasoning. By providing a challenging yet easy-to-evaluate suite, RIMO offers a high-resolution yardstick for future research, presenting a clear target for closing the profound reasoning gap our findings expose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。