arXiv:2602.06176cs.AIcs.CL2026-02被引 32

系统梳理大模型推理失败的三类根源,为提升智能可靠性提供指南。

Large Language Model Reasoning Failures

  • 按具身/非具身与直觉/逻辑区分推理类型
  • 发现三类失败:架构缺陷、领域局限、抗扰性差
  • 开源全链路资料库,助力研究者快速入门

大语言模型在多项任务中展现出卓越的推理能力,但仍在看似简单的场景中频繁出现推理失败。为系统理解并解决这些不足,本文首次提出针对大模型推理失败的综合性调研。我们构建了一个新的分类框架,将推理分为具身与非具身两类,其中非具身推理进一步细分为非正式(直觉)和正式(逻辑)推理。同时,从互补维度将推理失败划分为三类:一类是内在于大模型架构的根本性问题,广泛影响下游任务;第二类是特定应用领域的局限性表现;第三类是面对微小变化时性能不稳定的鲁棒性问题。针对每类失败,我们给出明确定义,分析已有研究,探究根本原因,并提出缓解策略。通过整合零散的研究工作,本调研为大模型推理中的系统性弱点提供了结构化视角,有助于未来研究向更强、更可靠、更鲁棒的推理能力演进。此外,我们还发布了涵盖相关研究的GitHub资源库(https://github.com/Peiyang-Song/Awesome-LLM-Reasoning-Failures),为该领域提供便捷入口。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have exhibited remarkable reasoning capabilities, achieving impressive results across a wide range of tasks. Despite these advances, significant reasoning failures persist, occurring even in seemingly simple scenarios. To systematically understand and address these shortcomings, we present the first comprehensive survey dedicated to reasoning failures in LLMs. We introduce a novel categorization framework that distinguishes reasoning into embodied and non-embodied types, with the latter further subdivided into informal (intuitive) and formal (logical) reasoning. In parallel, we classify reasoning failures along a complementary axis into three types: fundamental failures intrinsic to LLM architectures that broadly affect downstream tasks; application-specific limitations that manifest in particular domains; and robustness issues characterized by inconsistent performance across minor variations. For each reasoning failure, we provide a clear definition, analyze existing studies, explore root causes, and present mitigation strategies. By unifying fragmented research efforts, our survey provides a structured perspective on systemic weaknesses in LLM reasoning, offering valuable insights and guiding future research towards building stronger, more reliable, and robust reasoning capabilities. We additionally release a comprehensive collection of research works on LLM reasoning failures, as a GitHub repository at https://github.com/Peiyang-Song/Awesome-LLM-Reasoning-Failures, to provide an easy entry point to this area.

大模型推理失败系统综述可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。