测试大模型数学推理时,数字干扰会让其准确率下降超50%。
Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
- 通过插入无关语句和缺失关键指令,测试模型抗干扰能力。
- 含数字的干扰使小模型准确率下降近10%,最大降幅达51.55%。
- 模型依赖记忆模板而非逻辑推理,适合研究可信AI的读者看。
大语言模型在数学推理领域取得进展,但其是否具备真正的数学理解能力仍存争议。为此,我们提出一种新的扰动框架,通过注入语义无关的扰动句子并逐步增加强度,评估模型在复杂环境下的推理能力;同时引入核心提问指令缺失的扰动方法,深入分析模型求解机制。实验表明,面对无数字的扰动,模型表现相对稳定,但存在鲁棒性边界;当扰动含数字时,性能显著下降:多数开源小参数模型准确率下降近或超过10%,进一步增强扰动强度后,最大下降达51.55%;即使最先进的商业模型也出现3%-10%的性能损失。详细分析推理过程发现,模型对含数值的干扰更敏感,易受无关数字误导而给出错误答案,且干扰强度越高,缺陷越明显。此外,在缺少核心提问指令时,模型仍能保持20%-40%的准确率,表明其可能依赖记忆模板或模式匹配完成任务,而非真正逻辑推理。本研究揭示了当前大模型在推理能力上的局限性,对推动其发展具有重要意义。
原文摘要 · Abstract (English)
LLMs have made significant progress in the field of mathematical reasoning, but whether they have true the mathematical understanding ability is still controversial. To explore this issue, we propose a new perturbation framework to evaluate LLMs' reasoning ability in complex environments by injecting additional semantically irrelevant perturbation sentences and gradually increasing the perturbation intensity. At the same time, we use an additional perturbation method: core questioning instruction missing, to further analyze the LLMs' problem-solving mechanism. The experimental results show that LLMs perform stably when facing perturbation sentences without numbers, but there is also a robustness boundary. As the perturbation intensity increases, the performance exhibits varying degrees of decline; when facing perturbation sentences with numbers, the performance decreases more significantly, most open source models with smaller parameters decrease by nearly or even more than 10%, and further increasing with the enhancement of perturbation intensity, with the maximum decrease reaching 51.55%. Even the most advanced commercial LLMs have seen a 3%-10% performance drop. By analyzing the reasoning process of LLMs in detail, We find that models are more sensitive to perturbations with numerical information and are more likely to give incorrect answers when disturbed by irrelevant numerical information. The higher the perturbation intensity, the more obvious these defects are. At the same time, in the absence of core questioning instruction, models can still maintain an accuracy of 20%-40%, indicating that LLMs may rely on memory templates or pattern matching to complete the task, rather than logical reasoning. In general, our work reveals the shortcomings and limitations of current LLMs in their reasoning capabilities, which is of great significance for the further development of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。