自适应熵正则化提升大模型推理强化学习效果
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
- 动态调整熵正则系数,根据任务难度匹配探索强度
- 在多个数学推理数据集上准确率显著提升
- 适合需要稳定探索能力的LLM强化学习场景
推理能力已成为大语言模型的核心能力,基于可验证奖励的强化学习(RLVR)是提升该能力的关键范式。然而,RLVR训练常面临策略熵塌陷问题,即策略过于确定性,抑制探索并限制推理表现。尽管熵正则化是常见缓解手段,但其效果高度依赖固定系数,导致跨任务和模型表现不稳定。本文重新审视RLVR中的熵正则化,指出其潜力被低估。分析表明:(i) 不同难度任务需不同探索强度;(ii) 平衡探索需保持策略熵在初始值以下的适度范围。为此提出自适应熵正则化(AER)框架,包含三部分:难度感知系数分配、初始锚定目标熵和动态全局系数调整。多数学推理基准测试显示,AER持续优于基线,显著提升推理准确率与探索能力。
原文摘要 · Abstract (English)
Reasoning ability has become a defining capability of Large Language Models (LLMs), with Reinforcement Learning with Verifiable Rewards (RLVR) emerging as a key paradigm to enhance it. However, RLVR training often suffers from policy entropy collapse, where the policy becomes overly deterministic, hindering exploration and limiting reasoning performance. While entropy regularization is a common remedy, its effectiveness is highly sensitive to the fixed coefficient, making it unstable across tasks and models. In this work, we revisit entropy regularization in RLVR and argue that its potential has been largely underestimated. Our analysis shows that (i) tasks of varying difficulty demand distinct exploration intensities, and (ii) balanced exploration may require the policy entropy to be maintained within a moderate range below its initial level. Therefore, we propose Adaptive Entropy Regularization (AER)--a framework that dynamically balances exploration and exploitation via three components: difficulty-aware coefficient allocation, initial-anchored target entropy, and dynamic global coefficient adjustment. Experiments on multiple mathematical reasoning benchmarks show that AER consistently outperforms baselines, improving both reasoning accuracy and exploration capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。