提出新方法缓解大模型强化学习中的熵崩溃问题。
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
- 从熵变化角度分析强化学习中策略熵动态,建立统一理论框架。
- 发现现有方法仅调整部分因素,导致效果受限。
- 提出STEER方法,通过自适应重加权提升推理能力,性能优于基线。
基于可验证奖励的强化学习(RLVR)是提升大语言模型推理能力的核心技术,但其训练常受熵崩溃困扰——策略熵迅速下降,限制探索并降低训练效率。尽管已有研究尝试通过启发式熵干预缓解该问题,但机制尚不清晰。本文通过理论与实证分析,揭示了每步更新中词级熵变化的紧密解析近似,识别出四个决定性因素,并构建统一理论框架解释现有方法的影响。进一步发现:近期方法仅对其中一两个因素进行启发式调整,忽略其他关键因素,因而存在根本局限。基于此,提出STEER方法,依据理论估算的熵变化自适应重加权词元。在六个数学推理与三个编码基准上的实验表明,STEER能有效缓解熵崩溃,持续超越现有先进基线。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) serves as a cornerstone technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, its training is often plagued by \emph{entropy collapse}, a rapid decline in policy entropy that limits exploration and undermines training effectiveness. While recent works attempt to mitigate this issue via several heuristic entropy interventions, the underlying mechanisms remain poorly understood. In this work, we conduct comprehensive theoretical and empirical analyses of entropy dynamics in RLVR, offering two main insights: (1) We derive a tight analytical approximation for token-level entropy change at each update step, revealing four governing factors and providing a unified theoretical framework to explain how existing methods influence entropy; (2) We reveal a fundamental limitation of recent approaches: they rely on heuristic adjustments to one or two of these factors, leaving other relevant factors unconsidered, thus inherently limiting their effectiveness. Motivated by these findings, we propose STEER, a principled entropy-modulation method that adaptively reweights tokens based on theoretically-estimated entropy variations. Extensive experiments across six mathematical reasoning and three coding benchmarks demonstrate that STEER effectively mitigates entropy collapse and consistently outperforms state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。