通过可控实验揭示预训练、中段训练与强化学习对推理模型的真实作用
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- 构建合成推理任务框架,分离预训练、中段训练和强化学习的因果影响
- 强化学习仅在预训练留有提升空间时有效,且需针对模型能力边界设计任务
- 中段训练在固定算力下比纯强化学习更有效,过程奖励可减少奖励滥用
近期强化学习(RL)技术显著提升了语言模型的推理能力,但尚不清楚后训练是否真正扩展了模型在预训练阶段已获得的推理能力。核心挑战在于现代训练流程缺乏控制:大规模预训练语料不透明,中段训练常被忽视,而强化学习目标与未知先验知识复杂交互。为解决这一模糊性,我们构建了一个完全受控的实验框架,分离预训练、中段训练和基于强化学习的后训练的因果贡献。方法采用具有显式原子操作的合成推理任务,可解析的逐步推理轨迹,并系统操控训练分布。我们在两个维度评估模型:对更复杂组合的外推泛化能力,以及跨表面语境的上下文泛化能力。结果表明:1)强化学习仅在预训练留下足够提升空间,且强化学习数据针对模型能力边界(任务难但尚未超出范围)时,才能产生真实能力提升(pass@128);2)上下文泛化只需最小但足够的预训练暴露,之后强化学习可可靠迁移;3)在固定计算资源下,中段训练显著优于仅使用强化学习,凸显其在训练流程中的核心但被低估的作用;4)过程级奖励可减少奖励滥用,提升推理保真度。这些发现阐明了预训练、中段训练与强化学习之间的相互作用,为理解与改进推理语言模型的训练策略提供了基础。
原文摘要 · Abstract (English)
Recent reinforcement learning (RL) techniques have yielded impressive reasoning improvements in language models, yet it remains unclear whether post-training truly extends a model's reasoning ability beyond what it acquires during pre-training. A central challenge is the lack of control in modern training pipelines: large-scale pre-training corpora are opaque, mid-training is often underexamined, and RL objectives interact with unknown prior knowledge in complex ways. To resolve this ambiguity, we develop a fully controlled experimental framework that isolates the causal contributions of pre-training, mid-training, and RL-based post-training. Our approach employs synthetic reasoning tasks with explicit atomic operations, parseable step-by-step reasoning traces, and systematic manipulation of training distributions. We evaluate models along two axes: extrapolative generalization to more complex compositions and contextual generalization across surface contexts. Using this framework, we reconcile competing views on RL's effectiveness. We show that: 1) RL produces true capability gains (pass@128) only when pre-training leaves sufficient headroom and when RL data target the model's edge of competence, tasks at the boundary that are difficult but not yet out of reach. 2) Contextual generalization requires minimal yet sufficient pre-training exposure, after which RL can reliably transfer. 3) Mid-training significantly enhances performance under fixed compute compared with RL only, demonstrating its central but underexplored role in training pipelines. 4) Process-level rewards reduce reward hacking and improve reasoning fidelity. Together, these results clarify the interplay between pre-training, mid-training, and RL, offering a foundation for understanding and improving reasoning LM training strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。