arXiv:2507.07313cs.CLcs.AI2025-07被引 19

顶尖大模型在简单推理任务上仍频频失败,暴露其泛化能力短板。

Frontier LLMs Still Struggle with Simple Reasoning Tasks

  • 构建可调节难度的程序化推理任务,测试模型在复杂度上升时的表现。
  • 即使先进模型也因中间步骤错误、长上下文处理难而持续出错。
  • 对简化版经典谜题反而失效,暴露其对原始记忆的依赖问题。

尽管前沿大语言模型在数学与编程等挑战性基准上表现优异,却常在人类轻松完成的简单推理任务中失败。本文系统研究了这些‘简单’推理问题的表现,通过扩展已有工作,创建了一套可程序生成的简单推理任务,涵盖计数、一阶逻辑、证明树和旅行规划等,其参数(如文档长度或数学问题变量数)可任意调整以增加计算量,同时保持基础难度不变。此前研究显示传统非思考模型可被设计为失败,而本文表明,即使最先进的思考型模型也持续在这些任务中失败,原因类似(如统计捷径、中间步骤错误、长上下文处理困难)。我们进一步引入Unpuzzles数据集,一个由经典数学与逻辑谜题简化版组成的‘简单’基准。有趣的是,现代大模型虽擅长解决原版谜题,却常在简化版本中失败,表现出系统性错误模式,与对原始版本的记忆相关。即便模型能解决描述不同但逻辑相同的题目,此现象仍存在。结果表明,前沿语言模型在简单推理任务上的分布外泛化能力依然薄弱,任务变简单并不意味着性能提升。

原文摘要 · Abstract (English)

While state-of-the-art large language models (LLMs) demonstrate advanced reasoning capabilities-achieving remarkable performance on challenging competitive math and coding benchmarks-they also frequently fail on tasks that are easy for humans. This work studies the performance of frontier LLMs on a broad set of such "easy" reasoning problems. By extending previous work in the literature, we create a suite of procedurally generated simple reasoning tasks, including counting, first-order logic, proof trees, and travel planning, with changeable parameters (such as document length. or the number of variables in a math problem) that can arbitrarily increase the amount of computation required to produce the answer while preserving the fundamental difficulty. While previous work showed that traditional, non-thinking models can be made to fail on such problems, we demonstrate that even state-of-the-art thinking models consistently fail on such problems and for similar reasons (e.g. statistical shortcuts, errors in intermediate steps, and difficulties in processing long contexts). To further understand the behavior of the models, we introduce the unpuzzles dataset, a different "easy" benchmark consisting of trivialized versions of well-known math and logic puzzles. Interestingly, while modern LLMs excel at solving the original puzzles, they tend to fail on the trivialized versions, exhibiting several systematic failure patterns related to memorizing the originals. We show that this happens even if the models are otherwise able to solve problems with different descriptions but requiring the same logic. Our results highlight that out-of-distribution generalization is still problematic for frontier language models and the new generation of thinking models, even for simple reasoning tasks, and making tasks easier does not necessarily imply improved performance.

大模型推理能力泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。