让大模型像人一样慢思考,提升复杂任务推理能力。
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
- 用动态计算资源应对不同难度任务,实现智能调度
- 结合强化学习与自进化机制,持续优化决策过程
- 适合研究推理增强、AI决策系统的人参考
本综述探讨了近期旨在模拟人类'慢思考'的推理型大语言模型进展,其灵感源自卡尼曼《思考,快与慢》中的认知理论。这类模型如OpenAI的o1,在数学推理、视觉理解、医疗诊断及多智能体辩论等复杂任务中,通过动态扩展计算资源实现高效推理。基于超过100项研究的分析,本文梳理了三类关键技术:(1)测试时缩放,依据任务复杂度动态调整计算,包括搜索、采样与动态验证;(2)强化学习,通过策略网络、奖励模型与自我演化实现决策迭代优化;(3)慢思考框架(如长链思维、分层处理),将问题分解为可管理步骤。综述还指出了该领域面临的挑战与未来方向。提升大模型的推理能力对推动其在科学发现与决策支持等真实场景的应用至关重要。
原文摘要 · Abstract (English)
This survey explores recent advancements in reasoning large language models (LLMs) designed to mimic "slow thinking" - a reasoning process inspired by human cognition, as described in Kahneman's Thinking, Fast and Slow. These models, like OpenAI's o1, focus on scaling computational resources dynamically during complex tasks, such as math reasoning, visual reasoning, medical diagnosis, and multi-agent debates. We present the development of reasoning LLMs and list their key technologies. By synthesizing over 100 studies, it charts a path toward LLMs that combine human-like deep thinking with scalable efficiency for reasoning. The review breaks down methods into three categories: (1) test-time scaling dynamically adjusts computation based on task complexity via search and sampling, dynamic verification; (2) reinforced learning refines decision-making through iterative improvement leveraging policy networks, reward models, and self-evolution strategies; and (3) slow-thinking frameworks (e.g., long CoT, hierarchical processes) that structure problem-solving with manageable steps. The survey highlights the challenges and further directions of this domain. Understanding and advancing the reasoning abilities of LLMs is crucial for unlocking their full potential in real-world applications, from scientific discovery to decision support systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。