打造奥运级别数学难题集,测试大模型真实推理能力
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
- 构建350道奥数题双语库,分计算与形式化验证两类评估
- 模型在难题上表现远低于人类,暴露出靠猜测而非逻辑推导的缺陷
- 适合研究模型真实推理能力或跨语言对比的学者使用
大型推理模型的快速发展已使现有数学评测趋于饱和,亟需更具挑战性的评估框架。为此,我们推出OlymMATH——首个统一双范式评估的奥数级数学基准,包含350道题目,每题均有中英文双语版本。该基准涵盖两大评估模式:(1) 自然语言评估(OlymMATH-EASY与OlymMATH-HARD),含200道数值答案的计算题,支持客观规则评估;(2) 形式化验证(OlymMATH-LEAN),提供150道用Lean 4形式化的题目,实现过程级严格检验。所有题目均来自纸质出版物,人工采集以避免数据污染,经专家审核,覆盖四大核心领域。大量实验表明该基准具有显著挑战性,分析还揭示了语言间性能差异及模型依赖启发式‘猜测’而非严谨推理的现象。为支持社区研究,我们公开582,000+条推理轨迹、可视化工具及专家解法,详见https://github.com/RUCAIBox/OlymMATH。
原文摘要 · Abstract (English)
The rapid advancement of large reasoning models has saturated existing math benchmarks, underscoring the urgent need for more challenging evaluation frameworks. To address this, we introduce OlymMATH, a rigorously curated, Olympiad-level math benchmark comprising 350 problems, each with parallel English and Chinese versions. OlymMATH is the first benchmark to unify dual evaluation paradigms within a single suite: (1) natural language evaluation through OlymMATH-EASY and OlymMATH-HARD, comprising 200 computational problems with numerical answers for objective rule-based assessment, and (2) formal verification through OlymMATH-LEAN, offering 150 problems formalized in Lean 4 for rigorous process-level evaluation. All problems are manually sourced from printed publications to minimize data contamination, verified by experts, and span four core domains. Extensive experiments reveal the benchmark's significant challenge, and our analysis also uncovers consistent performance gaps between languages and identifies cases where models employ heuristic "guessing" rather than rigorous reasoning. To further support community research, we release 582k+ reasoning trajectories, a visualization tool, and expert solutions at https://github.com/RUCAIBox/OlymMATH.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。