揭示强化学习中熵与性能的动态权衡机制,提升大模型推理能力。
Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning
- 按熵变化将训练分为上升期与平台期,分层分析学习机制。
- 上升期降低负样本熵可加速有效推理模式学习,平台期高熵末尾词更易提升效率。
- 基于困惑度和位置动态调参,适配不同场景的强化学习优化。
近期,基于可验证奖励的强化学习(RLVR)被广泛用于增强大语言模型(LLMs)的推理能力。其核心挑战在于如何管理策略的熵与性能之间的权衡。尽管这一权衡至关重要,但对其在何种阶段、以何种方式最有效运作的理解仍不充分。为此,我们对RLVR中的熵-性能交换机制进行了系统性实证分析,涵盖阶段级、实例级和词元级三个粒度层次。结果表明,在上升阶段,负样本熵的降低有助于学习有效的推理模式,从而实现快速性能提升;在平台阶段,学习效率与低困惑度样本中高熵词元以及序列末端词元密切相关。基于此发现,我们提出两种方法,通过困惑度与位置信息动态调整奖励信号,聚焦于高学习潜力的词元,显著优于基线方法,在多个LLM上均取得改进。
原文摘要 · Abstract (English)
Recently, reinforcement learning with verifiable rewards (RLVR) has been widely used for enhancing the reasoning abilities of large language models (LLMs). A core challenge in RLVR involves managing the exchange between entropy and performance of policies. Despite the importance of this exchange, a fine-grained understanding of when and how this exchange operates most effectively remains limited. To bridge this gap, we conduct a systematic empirical analysis of the entropy-performance exchange mechanism of RLVR across different levels of granularity. Specifically, we first divide the training process into two distinct stages based on entropy dynamics, i.e., rising stage and plateau stage, and then systematically investigate how this mechanism varies across stage-level, instance-level, and token-level granularitiess. Our analysis reveals that, in the rising stage, entropy reduction in negative samples facilitates the learning of effective reasoning patterns, which in turn drives rapid performance gains. Moreover, in the plateau stage, learning efficiency strongly correlates with high-entropy tokens present in low-perplexity samples and those located at the end of sequences. Motivated by these findings, we propose two methods that dynamically adjust the reward signal using perplexity and positional information to focus RL updates on tokens that exhibit high learning potential, achieving improvements compared to the baseline methods on various LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。