揭示强化学习如何通过优化关键词重塑大模型推理能力
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
- 发现强化学习主要优化少数关键词,而非整体模式
- 不同推理模式成功率稳定,但训练中分布被重塑
- 适用于想理解强化学习提升推理机制的研究者
尽管强化学习在提升语言模型推理能力方面表现卓越,但其在大模型中的训练动态仍不清晰。本文通过系统性的推理模式级与词粒度分析,发现不同推理模式的成功率在训练中保持相对稳定,而强化学习主要优化一组稀疏的关键词,从而重塑推理模式分布并影响模型性能。基于这些实证发现,我们构建了理论框架,分析两类典型奖励机制下的训练动态:可验证奖励(RLVR)在两种特殊情况下分别呈现快速收敛与优化困难现象,表明基础模型的推理质量决定收敛行为;内部反馈奖励(RLIF)虽初期提升性能,但持续训练可能导致性能退化。大量实验验证了上述结论,推动了对强化学习在语言模型增强中理论与应用的理解。
原文摘要 · Abstract (English)
While reinforcement learning (RL) demonstrated remarkable success in enhancing the reasoning capabilities of language models, the training dynamics of RL in LLMs remain unclear. In this work, we provide an explanation of the RL training process through empirical analysis and rigorous theoretical modeling. First, through systematic reasoning-pattern-level and token-level analysis across the RL training process, we show that while different reasoning patterns exhibit relatively stable success rates during training, RL primarily optimizes a sparse subset of critical tokens, thereby reshaping reasoning pattern distributions to affect model performance. Building on these empirical insights, we develop a theoretical framework to understand the training dynamics of RL with two typical rewards: verifiable reward (RLVR) and model's internal feedback (RLIF). For RLVR, we analyze the training dynamics under two special cases: one where models readily converge to optimal reasoning strategies, and another where optimization becomes challenging, revealing that the base model's reasoning quality is crucial for determining convergence behavior. For RLIF, we examine how internal rewards initially improve model performance but can potentially lead to degradation with continued training. Extensive experiments validate our findings, advancing both theoretical understanding and practical applications of RL in language model enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。