构建面向奥数级数学推理的通用评测基准,挑战大模型极限。
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- 专为奥数级数学题设计,含4428道严格标注题目
- 顶尖模型在难题上准确率仅52.55%至60.54%
- 覆盖33个子领域与10种难度,适合评估高阶推理能力
大语言模型在数学推理方面取得显著进展,但现有基准如GSM8K或MATH已能被模型高精度解决(如OpenAI o1在MATH上达94.8%)。为填补这一差距,我们提出一个全面且具有挑战性的基准Omni-MATH,专门评估大模型在奥数级别的数学推理能力。不同于以往奥数类基准,本数据集仅聚焦数学,包含4428道竞赛级题目,经严格人工标注,细分为33个以上子领域,覆盖10余种难度等级,实现对模型在奥数推理中的全方位评估。实验表明,即使最先进模型OpenAI o1-mini和o1-preview在高难度奥数题上准确率也仅为60.54%和52.55%,凸显当前模型在奥数级数学推理上的显著不足。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8\% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54\% and 52.55\% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。