arXiv:2410.05229cs.LGcs.AI2024-10ICLR被引 636

用符号模板生成新题库,发现大模型数学推理极不稳。

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

  • 构建符号化题库GSM-Symbolic,可控生成多样化题目
  • 同一题改数字后模型表现普遍下降,最高降65%
  • 加无关条件就崩溃,说明模型不会真逻辑推理

大型语言模型(LLMs)在数学推理方面取得进展,常以GSM8K基准测试评估其能力。然而,性能提升是否真实仍存疑。为此,我们对多个最先进的开源与闭源模型进行了大规模研究。为克服现有评估局限,提出GSM-Symbolic——基于符号模板生成的改进型基准,支持更可控的评测。结果显示,所有模型在相同题型但数值变化时表现显著波动;当问题中增加无关但看似相关的语句时,性能急剧下降,最大降幅达65%,即便该语句不影响解题逻辑。这表明当前模型依赖训练数据中的模式复制,而非真正逻辑推理。本工作揭示了大模型在数学推理上的脆弱性与局限性。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have sparked interest in their formal reasoning capabilities, particularly in mathematics. The GSM8K benchmark is widely used to assess the mathematical reasoning of models on grade-school-level questions. While the performance of LLMs on GSM8K has significantly improved in recent years, it remains unclear whether their mathematical reasoning capabilities have genuinely advanced, raising questions about the reliability of the reported metrics. To address these concerns, we conduct a large-scale study on several SOTA open and closed models. To overcome the limitations of existing evaluations, we introduce GSM-Symbolic, an improved benchmark created from symbolic templates that allow for the generation of a diverse set of questions. GSM-Symbolic enables more controllable evaluations, providing key insights and more reliable metrics for measuring the reasoning capabilities of models.Our findings reveal that LLMs exhibit noticeable variance when responding to different instantiations of the same question. Specifically, the performance of all models declines when only the numerical values in the question are altered in the GSM-Symbolic benchmark. Furthermore, we investigate the fragility of mathematical reasoning in these models and show that their performance significantly deteriorates as the number of clauses in a question increases. We hypothesize that this decline is because current LLMs cannot perform genuine logical reasoning; they replicate reasoning steps from their training data. Adding a single clause that seems relevant to the question causes significant performance drops (up to 65%) across all state-of-the-art models, even though the clause doesn't contribute to the reasoning chain needed for the final answer. Overall, our work offers a more nuanced understanding of LLMs' capabilities and limitations in mathematical reasoning.

数学推理大模型测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。