arXiv:2509.08827cs.CLcs.AI2025-09综述被引 165

强化学习让大模型更会逻辑推理,这篇综述梳理了最新进展与挑战。

A Survey of Reinforcement Learning for Large Reasoning Models

论文配图:A Survey of Reinforcement Learning for Large Reasoning Models
图 1 · 摘自论文原文
  • 用强化学习提升大模型的数学与编程推理能力
  • 聚焦深求-1发布后的关键方法与训练资源进展
  • 适合关注大模型推理能力提升的研究者参考

本文综述了强化学习(RL)在大型语言模型(LLMs)推理能力提升中的最新进展。RL在解决数学、编程等复杂逻辑任务方面取得显著成果,已成为将LLMs转化为大推理模型(LRMs)的核心方法。随着领域快速发展,当前面临计算资源、算法设计、训练数据与基础设施等方面的可扩展性挑战。本文重点分析自深求-1发布以来,针对LLMs和LRMs的推理能力增强研究,涵盖基础组件、核心问题、训练资源及下游应用,旨在识别未来发展方向与机遇。希望推动更广泛的大模型推理研究。

原文摘要 · Abstract (English)

In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontier of LLM capabilities, particularly in addressing complex logical tasks such as mathematics and coding. As a result, RL has emerged as a foundational methodology for transforming LLMs into LRMs. With the rapid progress of the field, further scaling of RL for LRMs now faces foundational challenges not only in computational resources but also in algorithm design, training data, and infrastructure. To this end, it is timely to revisit the development of this domain, reassess its trajectory, and explore strategies to enhance the scalability of RL toward Artificial SuperIntelligence (ASI). In particular, we examine research applying RL to LLMs and LRMs for reasoning abilities, especially since the release of DeepSeek-R1, including foundational components, core problems, training resources, and downstream applications, to identify future opportunities and directions for this rapidly evolving area. We hope this review will promote future research on RL for broader reasoning models. Github: https://github.com/TsinghuaC3I/Awesome-RL-for-LRMs

强化学习大模型推理综述DeepSeek-R1

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。