用强化学习训练大模型分步思考,提升复杂推理能力。
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
- 用强化学习让大模型自动生成高质量推理路径。
- 测试时增加思考步骤可显著提高推理准确率。
- 适合研究大模型推理与智能系统构建的读者。
语言一直是人类推理的核心工具。大型语言模型(LLMs)的突破激发了研究者利用这些模型解决复杂推理任务的兴趣。研究者已超越简单的自回归生成,引入“思考”概念——即一系列代表推理中间步骤的标记序列。这一创新范式使LLM能够模拟树搜索和反思等复杂人类推理过程。近期,一种新兴趋势是通过强化学习(RL)训练模型掌握推理过程,借助试错搜索算法自动生成高质量推理轨迹,大幅扩展了LLM的推理能力,并提供大量训练数据。此外,研究表明,在测试时增加思考标记数可显著提升推理准确率。因此,训练时与测试时的规模化结合,开辟了通往大型推理模型的新前沿。OpenAI的o1系列标志着该方向的重要里程碑。本文综述了近期在LLM推理方面的进展,涵盖基础背景、关键核心技术(如自动化数据构建、学习推理方法、测试时扩展),分析主流开源项目,并总结开放挑战与未来方向。
原文摘要 · Abstract (English)
Language has long been conceived as an essential tool for human reasoning. The breakthrough of Large Language Models (LLMs) has sparked significant research interest in leveraging these models to tackle complex reasoning tasks. Researchers have moved beyond simple autoregressive token generation by introducing the concept of "thought" -- a sequence of tokens representing intermediate steps in the reasoning process. This innovative paradigm enables LLMs' to mimic complex human reasoning processes, such as tree search and reflective thinking. Recently, an emerging trend of learning to reason has applied reinforcement learning (RL) to train LLMs to master reasoning processes. This approach enables the automatic generation of high-quality reasoning trajectories through trial-and-error search algorithms, significantly expanding LLMs' reasoning capacity by providing substantially more training data. Furthermore, recent studies demonstrate that encouraging LLMs to "think" with more tokens during test-time inference can further significantly boost reasoning accuracy. Therefore, the train-time and test-time scaling combined to show a new research frontier -- a path toward Large Reasoning Model. The introduction of OpenAI's o1 series marks a significant milestone in this research direction. In this survey, we present a comprehensive review of recent progress in LLM reasoning. We begin by introducing the foundational background of LLMs and then explore the key technical components driving the development of large reasoning models, with a focus on automated data construction, learning-to-reason techniques, and test-time scaling. We also analyze popular open-source projects at building large reasoning models, and conclude with open challenges and future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。