arXiv:2501.12948cs.CLcs.AI2025-01被引 5.6k

用强化学习训练大模型,无需人工标注就能自发产生高级推理能力。

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

  • 纯强化学习激励大模型生成自我反思、验证等高级推理模式。
  • 在数学、编程等任务上超越依赖人工标注的监督学习模型。
  • 可提取推理模式用于提升小模型性能,适合想改进推理能力的研究者。

通用推理是人工智能领域长期存在的重大挑战。尽管大型语言模型(LLMs)与思维链提示已在基础推理任务中取得显著进展,但其成功高度依赖大量人工标注的推理轨迹,且对复杂问题仍能力不足。本文表明,通过纯粹的强化学习(RL),可有效激励大模型发展出高级推理能力,无需人类标注的推理路径。所提出的RL框架促进了自省、验证和动态策略调整等高级推理模式的涌现。训练后的模型在可验证任务(如数学、编程竞赛、科学、技术、工程与数学领域)中表现更优,超越了基于人类示范的监督学习模型。此外,这些大规模模型涌现出的推理模式可系统性地用于引导和增强小模型的推理能力。

原文摘要 · Abstract (English)

General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions, and STEM fields, surpassing its counterparts trained via conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models.

大模型强化学习推理能力自主生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。