揭示强化学习提升推理能力的两种内在机制。
Select and Improve: Understanding the Mechanics of Post-Training for Reasoning
- 通过策略选择和策略改进两类机制实现推理能力提升。
- 多样化推理策略监督可激活策略选择,难度递增训练数据促进策略改进。
- 为优化推理模型训练提供可操作的实证指导,适合模型开发者参考。
强化学习已成为推理与编程模型训练的关键组成部分,但其内在机制仍不清晰。本文基于对 Qwen-2.5-1.5B 模型的受控数学推理实验,揭示了能力提升的两大核心机制:策略选择与策略改进。研究发现,监督数据中包含多样化的推理策略可激活策略选择机制;而强化学习数据的难度逐步提升则能驱动策略改进。这些结果从机制层面深化了对强化学习训练的理解,并提出了可实际应用的干预方法,有助于持续扩展模型的推理能力。
原文摘要 · Abstract (English)
Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We study how and through what underlying processes capabilities are acquired or enhanced via reinforcement learning post-training. Our analysis, based on controlled math reasoning experiments with Qwen-2.5-1.5B, reveals two core mechanisms: strategy selection and strategy improvement. Our results highlight the role of SFT data and reinforcement learning data in activating these mechanisms, in particular showing how supervising the model on diverse reasoning strategies can enable strategy selection and how increasing difficulty in reinforcement learning data can enable strategy improvement. Taken together, our results provide mechanistic insight into RL training and suggest practical interventions to continue scaling reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。