测试大模型在数学难题扰动下的推理能力,发现真实推理能力有限。
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
- 构建了两类扰动题集:简单扰动和根本性改变的硬扰动。
- 模型在硬扰动题上性能下降显著,如o1-mini降16.49%。
- 提醒警惕机械套用解法的伪推理,适合关注模型鲁棒性的研究者。
大型语言模型在复杂数学推理任务中表现优异,引发其是否具备真正推理能力还是仅靠记忆的讨论。以往研究通过简单扰动(保持原解题逻辑)构建基准,但尚未探索根本性改变问题本质的硬扰动。为此,我们基于MATH数据集中最难的Level-5题目,分别构造了MATH-P-Simple与MATH-P-Hard两个题集,各包含279道扰动题。实验发现,多种模型在MATH-P-Hard上性能大幅下降,如o1-mini下降16.49%,gemini-2.0-flash-thinking下降12.9%。我们还指出一种新型记忆现象:模型盲目套用习得解法,不评估其在新情境下的适用性,尤其在使用原始题进行上下文学习时更严重。该问题对构建更可靠、鲁棒的推理模型至关重要,亟需后续研究应对。
原文摘要 · Abstract (English)
Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieved by true reasoning capability or memorization. To investigate this question, prior work has constructed mathematical benchmarks when questions undergo simple perturbations -- modifications that still preserve the underlying reasoning patterns of the solutions. However, no work has explored hard perturbations, which fundamentally change the nature of the problem so that the original solution steps do not apply. To bridge the gap, we construct MATH-P-Simple and MATH-P-Hard via simple perturbation and hard perturbation, respectively. Each consists of 279 perturbed math problems derived from level-5 (hardest) problems in the MATH dataset (Hendrycksmath et. al., 2021). We observe significant performance drops on MATH-P-Hard across various models, including o1-mini (-16.49%) and gemini-2.0-flash-thinking (-12.9%). We also raise concerns about a novel form of memorization where models blindly apply learned problem-solving skills without assessing their applicability to modified contexts. This issue is amplified when using original problems for in-context learning. We call for research efforts to address this challenge, which is critical for developing more robust and reliable reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。