小样本蒸馏比零样本强化学习更擅长培养灵活推理能力
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
- 仅用920个样本的蒸馏方法提升模型推理灵活性
- 蒸馏模型使用更多拟人化词汇和逻辑连接词
- 适合关注高效推理增强的研究者或工程师
强化学习(RL)在提升大语言模型(LLMs)推理能力方面发挥重要作用。一些研究直接对较小的基础模型应用RL(称为零样本强化学习,zero-RL),也取得了显著进展。然而本文表明,仅使用920个示例,基于基础模型的简单蒸馏方法即可明显优于通常需要大量数据和计算成本的零样本强化学习。通过分析模型输出中的词元频率,发现蒸馏模型表现出更强的灵活推理能力,其使用拟人化词元和逻辑连接词的频率远高于零样本强化学习模型。进一步分析显示,蒸馏增强了两种高级认知行为:多视角思考/尝试和元认知意识。这些行为的频繁出现促成了灵活推理,而零样本强化学习未能显著提升此类行为的频率。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has played an important role in improving the reasoning ability of large language models (LLMs). Some studies apply RL directly to \textit{smaller} base models (known as zero-RL) and also achieve notable progress. However, in this paper, we show that using only 920 examples, a simple distillation method based on the base model can clearly outperform zero-RL, which typically requires much more data and computational cost. By analyzing the token frequency in model outputs, we find that the distilled model shows more flexible reasoning. It uses anthropomorphic tokens and logical connectors much more often than the zero-RL model. Further analysis reveals that distillation enhances the presence of two advanced cognitive behaviors: Multi-Perspective Thinking or Attempting and Metacognitive Awareness. Frequent occurrences of these two advanced cognitive behaviors give rise to flexible reasoning, which is essential for solving complex reasoning problems, while zero-RL fails to significantly boost the frequency of these behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。