arXiv:2606.11470cs.CL2026-06综述

梳理大模型推理的9类范式与常见失败模式,提供系统性参考框架。

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

论文配图:The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes
图 1 · 摘自论文原文
  • 按推理范式、方法机制和失败模式分类300+论文,构建系统框架。
  • 揭示幻觉推理、多步脆弱、虚假理由等共性失败问题。
  • 适合研究者评估模型能力或设计更鲁棒的推理系统。

推理已成为大语言模型(LLMs)评估与解释的核心,涵盖思维链(CoT)、数学求解、多跳问答、代码生成、检索增强推理、工具使用及多模态决策等多个方向。本文提出「大模型推理周期表」,对300余篇近期论文按推理范式、方法机制、评估设置和失败模式进行组织。将LLM推理分为九类:思维链、多跳、数学、常识、视觉与时间、代码与算法、检索增强、工具增强或代理型、基于强化学习的推理。针对每类,综述提示工程、架构干预、监督微调、验证器引导推理、奖励建模、检索、工具接口、代理工作流与基准设计等方法。指出推理并非单一涌现能力,而是由模型规模、任务结构、外部记忆、监督信号与评估协议共同塑造的系列行为。总结出反复出现的失败模式,包括幻觉推理、脆弱的多步推演、虚假推理、弱因果基础、泛化能力差、基准污染与不可靠自验证。跨范式进展难以比较,因性能提升可能源于提示、检索、验证器设计或基准结构,而非通用推理能力。该综述连接方法与其假设、优势与缺陷,为领域提供参考地图与未来研究诊断框架。结论认为,鲁棒推理需元推理、多模态与时间对齐、自适应工具使用及分布外变化下的规范评估。

原文摘要 · Abstract (English)

Reasoning has become central to how Large Language Models (LLMs) are evaluated and interpreted, spanning Chain-of-Thought (CoT), mathematical problem-solving, multi-hop question answering, code generation, retrieval-augmented reasoning, tool use, and multimodal decision-making. In this survey, we introduce the Periodic Table of LLM Reasoning, a framework organizing 300+ recent papers by reasoning paradigm, methodological mechanism, evaluation setting, and failure mode. We classify LLM reasoning into nine paradigms: Chain-of-Thought, Multi-Hop, Mathematical, Commonsense, Visual and Temporal, Code and Algorithmic, Retrieval-Augmented, Tool-Augmented or Agentic, and Reinforcement Learning-based reasoning. For each, we review approaches, including prompting, architectural interventions, supervised fine-tuning, verifier-guided inference, reward modeling, retrieval, tool interfaces, agentic workflows, and benchmark design. We argue that LLM reasoning is not a single emergent capability but a family of scaffolded behaviors shaped by model scale, task structure, external memory, supervision, and evaluation protocols. We synthesize recurring failure modes, including hallucinated reasoning, brittle multi-step inference, spurious rationales, weak causal grounding, poor out-of-distribution generalization, benchmark contamination, and unreliable self-verification. Progress is difficult to compare across paradigms because gains may arise from prompting, retrieval, verifier design, or benchmark structure rather than general reasoning ability. The survey connects methods to their assumptions, strengths, and failure modes, providing a reference map of the field and a diagnostic framework for future work. We conclude that robust LLM reasoning will require meta-reasoning, multimodal and temporal grounding, adaptive tool use, and principled evaluation under distribution shift.

大模型推理周期表框架失败模式综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。