用高质量推理模板初始化,让大模型强化学习更高效
Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
- 通过自动发现和整理推理模板,丰富初始知识
- 在数学逻辑任务上,显著提升强化学习初期表现
- 适合希望加速大模型推理训练的研究者
强化学习(RL)已成为提升大语言模型(LLMs)推理能力的重要方法。尽管取得显著进展,仍面临采样效率低、对模型初始化高度敏感等挑战:部分模型仅需少量强化学习步骤即可快速提升,而另一些则需大量数据才能见效。本文从推理令牌覆盖率角度出发,提出初始化时引入多样且高质量的推理原语是实现稳定、高效强化学习的关键。为此,我们设计了Tailor微调流程,能自动发现并筛选新的推理原语,从而在强化学习前扩大推理状态分布的覆盖范围。在数学与逻辑推理基准上的大量实验表明,Tailor生成的预热数据更具多样性与质量,显著提升了下游强化学习性能。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs). While RL has demonstrated substantial performance gains, it still faces key challenges, including low sampling efficiency and a strong dependence on model initialization: some models achieve rapid improvements with minimal RL steps, while others require significant training data to make progress. In this work, we investigate these challenges through the lens of reasoning token coverage and argue that initializing LLMs with diverse, high-quality reasoning primitives is essential for achieving stable and sample-efficient RL training. We propose Tailor, a finetuning pipeline that automatically discovers and curates novel reasoning primitives, thereby expanding the coverage of reasoning-state distributions before RL. Extensive experiments on mathematical and logical reasoning benchmarks demonstrate that Tailor generates more diverse and higher-quality warm-start data, resulting in higher downstream RL performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。