通过梯度保持剪裁重设熵控制,缓解大模型强化学习中的过早自信问题。
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
- 从梯度保持剪裁角度重构熵控制机制,揭示重要性采样区域对熵的影响。
- 设计动态剪裁阈值与多种熵变化策略,有效抑制熵塌陷。
- 适用于需提升推理多样性与稳定性的大语言模型强化学习场景。
基于可验证奖励的强化学习(RLVR)已成为提升大语言模型推理能力的关键方法。然而,持续训练常导致策略熵塌陷,表现为熵快速下降,引发过早自信、输出多样性降低以及梯度范数消失,阻碍学习进程。梯度保持剪裁是影响该现象的主要因素,但现有缓解策略多为静态且缺乏将剪裁机制与精确熵控制相联系的框架。本文从梯度保持剪裁视角重塑熵控制,理论与实证验证了特定重要性采样比区域对熵增长与减少的贡献。基于此,提出使用动态剪裁阈值的新型调节机制以精准管理熵。进一步设计并评估了多种动态熵控制策略,包括先增后减、先减后增再减及振荡衰减。实验结果表明,这些策略能有效缓解熵塌陷,并在多个基准测试中实现更优性能。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a critical method for enhancing the reasoning capabilities of Large Language Models (LLMs). However, continuous training often leads to policy entropy collapse, characterized by a rapid decay in entropy that results in premature overconfidence, reduced output diversity, and vanishing gradient norms that inhibit learning. Gradient-Preserving Clipping is a primary factor influencing these dynamics, but existing mitigation strategies are largely static and lack a framework connecting clipping mechanisms to precise entropy control. This paper proposes reshaping entropy control in RL from the perspective of Gradient-Preserving Clipping. We first theoretically and empirically verify the contributions of specific importance sampling ratio regions to entropy growth and reduction. Leveraging these findings, we introduce a novel regulation mechanism using dynamic clipping thresholds to precisely manage entropy. Furthermore, we design and evaluate dynamic entropy control strategies, including increase-then-decrease, decrease-increase-decrease, and oscillatory decay. Experimental results demonstrate that these strategies effectively mitigate entropy collapse and achieve superior performance across multiple benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。