arXiv:2501.13766cs.CLcs.AI2025-01ICLR被引 16

构建首个面向本科生数学推理的动态评估基准,揭示大模型在题目变化下的脆弱性。

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

  • 设计包含5062道题的多主题动态基准,每题三版本防测试污染
  • 提出有效准确率与推理差距新指标,发现最高准确率仅56.3%
  • 适合关注数学推理鲁棒性的研究者和模型开发者使用

大型语言模型在数学推理方面已取得显著进展,亟需全面且公平的评估体系。现有基准普遍存在覆盖范围不足或测试集污染问题。为此,我们提出UGMathBench,一个专为评估本科生数学推理能力设计的多样化、动态化基准。该基准涵盖16个学科、111个主题,共5,062道题目,包含10种不同答案类型。每道题有三个随机版本,未来将根据主流开源模型饱和情况持续补充新版本。我们提出两个关键指标:有效准确率(EAcc),衡量所有版本中正确解题的比例;推理差距(Δ),通过计算各版本平均准确率与EAcc的差值,评估推理鲁棒性。对23个领先大模型的评估显示,最高EAcc仅为56.3%(OpenAI-o1-mini),且各模型均呈现较大Δ值。这表明亟需发展具备高有效准确率且Δ=0的‘大推理模型’。我们预计,UGMathBench及其详细评估代码的发布,将有力推动大模型在数学求解方向的发展。代码与数据见:https://github.com/YangLabHKUST/UGMathBench

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made significant strides in mathematical reasoning, underscoring the need for a comprehensive and fair evaluation of their capabilities. However, existing benchmarks often fall short, either lacking extensive coverage of undergraduate-level mathematical problems or probably suffering from test-set contamination. To address these issues, we introduce UGMathBench, a diverse and dynamic benchmark specifically designed for evaluating undergraduate-level mathematical reasoning with LLMs. UGMathBench comprises 5,062 problems across 16 subjects and 111 topics, featuring 10 distinct answer types. Each problem includes three randomized versions, with additional versions planned for release as leading open-source LLMs become saturated in UGMathBench. Furthermore, we propose two key metrics: effective accuracy (EAcc), which measures the percentage of correctly solved problems across all three versions, and reasoning gap ($Δ$), which assesses reasoning robustness by calculating the difference between the average accuracy across all versions and EAcc. Our extensive evaluation of 23 leading LLMs reveals that the highest EAcc achieved is 56.3\% by OpenAI-o1-mini, with large $Δ$ values observed across different models. This highlights the need for future research aimed at developing "large reasoning models" with high EAcc and $Δ= 0$. We anticipate that the release of UGMathBench, along with its detailed evaluation codes, will serve as a valuable resource to advance the development of LLMs in solving mathematical problems. Codes and data are available at https://github.com/YangLabHKUST/UGMathBench

数学推理评估基准大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。