arXiv:2412.09413cs.AIcs.CL2024-12被引 163

复现类o1慢思考模型,通过模仿、探索与自迭代提升推理能力

Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

  • 采用模仿-探索-自改进框架,分阶段训练推理模型
  • 在三个基准上表现媲美工业级推理系统
  • 适合研究推理机制与可复现性的人参考

近期,如o1的慢思考推理系统在解决复杂推理任务上展现出显著能力。这类系统通过延长思考过程,生成更全面、准确且有逻辑的答案。然而,其核心技术多由产业界掌握,未公开披露。为此,学术界正积极探究其技术基础。本文基于前期工作,报告了实现类o1推理系统的复现研究。提出名为STILL-2的“模仿、探索、自改进”框架作为核心方法:第一阶段使用提炼的长思维链数据微调模型,使其具备慢思考模式;第二阶段通过生成多个回溯轨迹,引导模型探索难题,逐步积累高质量解题路径;第三阶段通过迭代优化训练数据实现模型自改进。我们在三个挑战性基准上进行大量实验,结果表明该方法在性能上达到与工业级推理系统相当的水平。

原文摘要 · Abstract (English)

Recently, slow-thinking reasoning systems, such as o1, have demonstrated remarkable capabilities in solving complex reasoning tasks. These systems typically engage in an extended thinking process before responding to a query, allowing them to generate more thorough, accurate, and well-reasoned solutions. These systems are primarily developed and maintained by industry, with their core techniques not publicly disclosed. In response, an increasing number of studies from the research community aim to explore the technical foundations underlying these powerful reasoning systems. Building on these prior efforts, this paper presents a reproduction report on implementing o1-like reasoning systems. We introduce an ``imitate, explore, and self-improve'' framework, denoted as \textbf{STILL-2}, as our primary technical approach to train the reasoning model. In the initial phase, we use distilled long-form thought data to fine-tune the reasoning model, enabling it to invoke a slow-thinking mode. The model is then encouraged to explore challenging problems by generating multiple rollouts, which can result in increasingly more high-quality trajectories that lead to correct answers. Furthermore, the model undergoes self-improvement by iteratively refining its training dataset. To verify the effectiveness of this approach, we conduct extensive experiments on three challenging benchmarks. The experimental results demonstrate that our approach achieves competitive performance compared to industry-level reasoning systems on these benchmarks.

推理系统慢思考自改进复现研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。