复现DeepSeek-R1的进展与未来方向综述
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
- 梳理主流复现方法,聚焦监督微调与可验证奖励强化学习
- 多个团队实现接近原模型性能,依赖开源数据与训练流程
- 适合关注推理模型优化与开源复现的研究者参考
近期推理语言模型(RLMs)的发展代表了大模型的一次新演进。特别是DeepSeek-R1的发布引发了广泛关注,并激发了研究界对语言模型显式推理范式的探索热情。然而,DeepSeek未完全开源其发布模型的实现细节,包括DeepSeek-R1-Zero、DeepSeek-R1及蒸馏小型模型。因此,大量复现研究涌现,旨在重现DeepSeek-R1的优异表现,通过相似训练流程和完全开源的数据资源,实现了性能相近的结果。这些工作集中于监督微调(SFT)和基于可验证奖励的强化学习(RLVR),在数据准备与方法设计方面提出了多种可行策略,积累了宝贵经验。本文总结了近期复现研究,重点介绍SFT与RLVR的细节,涵盖数据构建、方法设计与训练流程。同时归纳了各研究中实现细节与实验结果的关键发现,以期启发未来研究。此外,还讨论了增强推理模型的其他技术,强调扩展应用前景并分析发展挑战。本综述旨在帮助研究者与开发者掌握最新进展,激发进一步提升推理语言模型的新思路。
原文摘要 · Abstract (English)
The recent development of reasoning language models (RLMs) represents a novel evolution in large language models. In particular, the recent release of DeepSeek-R1 has generated widespread social impact and sparked enthusiasm in the research community for exploring the explicit reasoning paradigm of language models. However, the implementation details of the released models have not been fully open-sourced by DeepSeek, including DeepSeek-R1-Zero, DeepSeek-R1, and the distilled small models. As a result, many replication studies have emerged aiming to reproduce the strong performance achieved by DeepSeek-R1, reaching comparable performance through similar training procedures and fully open-source data resources. These works have investigated feasible strategies for supervised fine-tuning (SFT) and reinforcement learning from verifiable rewards (RLVR), focusing on data preparation and method design, yielding various valuable insights. In this report, we provide a summary of recent replication studies to inspire future research. We primarily focus on SFT and RLVR as two main directions, introducing the details for data construction, method design and training procedure of current replication studies. Moreover, we conclude key findings from the implementation details and experimental results reported by these studies, anticipating to inspire future research. We also discuss additional techniques of enhancing RLMs, highlighting the potential of expanding the application scope of these models, and discussing the challenges in development. By this survey, we aim to help researchers and developers of RLMs stay updated with the latest advancements, and seek to inspire new ideas to further enhance RLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。