梳理大模型从直觉到深度推理的演进路径,解析专家级推理能力的实现机制。
From System 1 to System 2: A Survey of Reasoning Large Language Models

- 对比系统1(快速直觉)与系统2(逐步推理)的认知差异,定位大模型短板。
- 总结o1/o3、R1等模型在数学与编程任务中逼近人类专家的表现。
- 适合关注大模型推理能力、认知模拟的研究者与开发者参考。
实现类人智能需要优化从快速直觉的系统1向更慢但更精准的系统2推理的过渡。基础大语言模型擅长快速决策,但在复杂推理上缺乏深度,尚未充分具备系统2式的逐步分析能力。近期,如OpenAI的o1/o3和DeepSeek的R1等推理型大模型在数学与编程领域展现出专家级表现,接近人类的深思熟虑过程,体现出类人认知能力。本文首先回顾基础大模型的发展及早期系统2技术的探索,阐明二者融合如何推动推理型大模型的诞生。随后分析推理型大模型的构建方法、核心机制及其演进路径,并对代表性模型在推理基准上的表现进行深入比较。最后探讨未来发展方向,并维护一个实时更新的GitHub仓库以追踪最新进展。本综述旨在为该快速发展的领域提供有价值参考,激发创新与进步。
原文摘要 · Abstract (English)
Achieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical reasoning for more accurate judgments and reduced biases. Foundational Large Language Models (LLMs) excel at fast decision-making but lack the depth for complex reasoning, as they have not yet fully embraced the step-by-step analysis characteristic of true System 2 thinking. Recently, reasoning LLMs like OpenAI's o1/o3 and DeepSeek's R1 have demonstrated expert-level performance in fields such as mathematics and coding, closely mimicking the deliberate reasoning of System 2 and showcasing human-like cognitive abilities. This survey begins with a brief overview of the progress in foundational LLMs and the early development of System 2 technologies, exploring how their combination has paved the way for reasoning LLMs. Next, we discuss how to construct reasoning LLMs, analyzing their features, the core methods enabling advanced reasoning, and the evolution of various reasoning LLMs. Additionally, we provide an overview of reasoning benchmarks, offering an in-depth comparison of the performance of representative reasoning LLMs. Finally, we explore promising directions for advancing reasoning LLMs and maintain a real-time \href{https://github.com/zzli2022/Awesome-Slow-Reason-System}{GitHub Repository} to track the latest developments. We hope this survey will serve as a valuable resource to inspire innovation and drive progress in this rapidly evolving field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。