用进化测试生成难题,让大模型数学评测更真实。
EvolMathEval: Towards Evolvable Benchmarks for Mathematical Reasoning via Evolutionary Testing
- 通过进化测试自动迭代生成高难度数学题。
- 使GSM8K数据集准确率平均下降48%。
- 发现模型常靠模糊条件假突破,适合评测者关注。
大型语言模型的快速进步使得现有数学推理基准逐渐变得简单,因为模型可从公开基准中学习。为应对这一挑战,本文提出EvolMathEval,一种基于进化测试的自动化数学基准生成与演化框架。实验表明,EvolMathEval不仅能持续生成大量高难度问题,还能显著提升GSM8K等公开数据集的复杂性,使模型平均准确率下降48%。深入分析发现,面对这些演化后的问题,大模型倾向于绕过复杂的多步逻辑推理,依赖简单模糊的条件得出错误答案,我们称此现象为“伪顿悟”,其在目标问题中导致了77%至100%的错误。代码与资源已公开。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) poses a significant challenge to existing mathematical reasoning benchmarks. However, these benchmarks tend to become easier over time as LLMs can learn from the published benchmarks. This limitation hinder the precise evaluation of the true capabilities of SOTA models. To address this challenge, this paper introduces EvolMathEval, an automated mathematical benchmark generation and evolution framework based on evolutionary testing. Experimental results demonstrate that EvolMathEval can not only generate a large volume of high-difficulty problems through continuous self-iteration, but it can also significantly enhance the complexity of public datasets like GSM8K through evolution, reducing model accuracy by an average of 48\%. Deeper investigation reveals that when solving these evolved problems, LLMs tend to bypass complex multi-step logical reasoning by relying on simplistic and fuzzy conditions, consequently leading to incorrect solutions. We define this phenomenon as the ``Pseudo Aha Moment", which we find accounts for 77\% to 100\% of errors on targeted problems. Code and resources are available at: https://anonymous.4open.science/r/EvolMathEval
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。