arXiv:2601.04670cs.LG2026-01被引 3

揭示强化学习微调中输出多样性下降的根源并提出新训练策略

Learning Dynamics in RL Post-Training for Language Models

  • 通过神经正切核分解分析更新传播机制
  • 发现特征表示变化有限导致模型置信度上升,引发多样性下降
  • 提出先优化分类器的两阶段策略,加速收敛且提升性能

强化学习(RL)后训练是现代语言模型开发的关键阶段,对提升对齐性和推理能力至关重要。然而,输出多样性降低等现象仍缺乏理解。本文从监督学习中已研究但RL中未充分探索的角度,采用经验性神经正切核(NTK)框架,将NTK分解为两个分量,以刻画RL更新在训练样本间的传播方式。分析表明,特征表示变化有限会导致RL更新系统性提升模型置信度,从而解释了后训练中输出多样性下降的现象。此外,我们发现该阶段的有效学习依赖于快速塑造分类器,这直接影响了NTK的梯度成分。基于此,我们提出分类器优先强化学习(CF-RL),一种先集中优化分类器再进行标准RL的两阶段策略。实验验证了理论分析,显示CF-RL下模型置信度提升且优化加速。进一步分析表明,CF-RL的作用机制不同于监督学习中的线性探测再微调。本研究形式化了RL后训练的学习动态,推动后续分析与改进。

原文摘要 · Abstract (English)

Reinforcement learning (RL) post-training is a critical stage in modern language model development, playing a key role in improving alignment and reasoning ability. However, several phenomena remain poorly understood, including the reduction in output diversity. To gain a broader understanding of RL post-training, we analyze the learning dynamics of RL post-training from a perspective that has been studied in supervised learning but remains underexplored in RL. We adopt an empirical neural tangent kernel (NTK) framework and decompose the NTK into two components to characterize how RL updates propagate across training samples. Our analysis reveals that limited variability in feature representations can cause RL updates to systematically increase model confidence, providing an explanation for the commonly observed reduction in output diversity after RL post-training. Furthermore, we show that effective learning in this regime depends on rapidly shaping the classifier, which directly affects the gradient component of the NTK. Motivated by these insights, we propose classifier-first reinforcement learning (CF-RL), a simple two-stage training strategy that prioritizes classifier updates before standard RL optimization. Experimental results validate our theoretical analysis by demonstrating increased model confidence and accelerated optimization under CF-RL. Additional analysis shows that the mechanism underlying CF-RL differs from that of linear-probing-then-fine-tuning in supervised learning. Overall, our study formalizes the learning dynamics of RL post-training and motivates further analysis and improvement.

强化学习语言模型学习动态模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。