arXiv:2502.08680cs.LGcs.AI2025-02被引 26

测试大模型在不同数值范围下的数学推理能力,发现其逻辑错误随数值复杂度显著上升。

Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges

  • 基于GSM8K生成新数据集,系统改变数值范围以评估模型鲁棒性。
  • 提出区分逻辑与非逻辑错误的新评分方法,揭示真实推理缺陷。
  • 模型在嵌入问题中的计算表现远差于独立算术任务,适合研究者参考。

大型语言模型的数学推理常在数值范围有限的基准上评估,无法反映真实世界跨尺度问题求解。现有方法仅对比输出与标准答案,难以洞察推理过程。为此,我们提出GSM-Ranges,基于GSM8K系统扰动数学题中数值,评估模型在不同数量级下的鲁棒性。同时,设计一种新评分机制,区分逻辑错误与非逻辑错误,更精准衡量推理质量。实验显示,随着数值复杂度增加,逻辑错误率最高提升14个百分点,表明模型对分布外数值存在普遍弱点。尽管模型在独立算术任务中准确率高,但在包含计算的文本问题中表现显著下降。该研究全面评估了大模型的数学推理能力,为提升数值泛化能力提供方向。

原文摘要 · Abstract (English)

Mathematical reasoning in Large Language Models (LLMs) is often evaluated using benchmarks with limited numerical ranges, failing to reflect real-world problem-solving across diverse scales. Furthermore, most existing evaluation methods only compare model outputs to ground-truth answers, obscuring insights into reasoning processes. To address these limitations, we introduce GSM-Ranges, a dataset generator derived from GSM8K that systematically perturbs numerical values in math problems to assess model robustness across varying numerical scales. Additionally, we propose a novel grading methodology that distinguishes between logical and non-logical errors, offering a more precise evaluation of reasoning processes beyond computational accuracy. Our experiments with various models reveal a significant increase in logical error rates-up to 14 percentage points-as numerical complexity rises, demonstrating a general weakness in reasoning with out-of-distribution numerical values. Moreover, while models demonstrate high accuracy on standalone arithmetic tasks, their performance deteriorates substantially when computations are embedded within word problems. These findings provide a comprehensive evaluation of LLMs' mathematical reasoning capabilities and inform future research directions for improving numerical generalization in language models.

数学推理大模型评估逻辑错误数值泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。