arXiv:2510.01227cs.CLcs.LG2025-10被引 1

新数学奥赛基准测试揭示大模型真实推理能力短板

EEFSUVA: A New Mathematical Olympiad Benchmark

  • 从东欧及前苏联国家竞赛中构建全新奥数题库
  • 顶尖大模型在新基准上表现显著下滑
  • 适合评估模型真实数学推理能力的科研人员

近期突破引发大语言模型(LLMs)在数学基准上达到国际数学奥林匹克(IMO)金牌水平的宣称。本文深入检验这些说法,评估现有基准是否真正反映模型的数学推理能力。当前基准多源于IMO及其相关赛事,可能因数据泄露和题型单一而高估模型表现。为此,我们引入EEFSUVA——一个源自东欧及前苏联地区区域性、国家级奥数竞赛的新基准。这些竞赛题目难度与IMO相当,以非标准解题技巧著称,但在网络语料中极为罕见。初步结果表明,即使最先进的LLMs在EEFSUVA上的表现也明显低于其他奥数类基准。这提示更广泛的评估数据集对全面评估数学推理能力及指导未来模型发展具有重要意义。

原文摘要 · Abstract (English)

Recent breakthroughs have spurred claims that large language models (LLMs) match gold medal Olympiad to graduate level proficiency on mathematics benchmarks. In this work, we examine these claims in detail and assess the extent to which current benchmarks capture genuine LLM mathematical reasoning. The composition of these benchmarks, primarily drawing from the International Mathematics Olympiad (IMO) and related competitions, may overstate models reasoning ability due to potential data contamination and a narrow focus on familiar problem types. To enable a more holistic assessment of mathematical understanding, we introduce EEFSUVA, a novel benchmark curated from under circulated regional and national Olympiads of Eastern Europe and the countries from the former Soviet Union. These contests feature problems of comparable difficulty to the IMO and are renowned for demanding nonstandard problem-solving techniques, yet their problems are far less prevalent in online corpora. Preliminary results suggest that even state-of-the-art LLMs exhibit a notable performance decline on EEFSUVA relative to other Olympiad-style benchmarks. These findings also suggest the potential importance of broader evaluation datasets for a fuller assessment of mathematical reasoning and for guiding future model development.

数学推理大模型评估奥数基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。