arXiv:2604.09676cs.LGcs.AI2026-04

对比两种熵控制方法,揭示为何新方法能避免过早收敛。

A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning

  • 用协方差机制选择性正则化高相关词元,减少偏差。
  • 传统方法引入持续偏移,导致策略次优;新方法可渐进无偏。
  • 适合做大模型推理训练与强化学习调优的研究者参考。

强化学习(RL)已成为提升大语言模型(LLMs)推理能力的关键方法,但规模化训练常受策略熵快速坍塌的阻碍,导致过早收敛和性能饱和。本文对两种熵控制策略进行理论比较:传统熵正则化与近期提出的基于协方差的机制。我们建立了一个统一框架,分析软最大参数化下的熵动态,发现熵变化由对数概率与逻辑值更新间的协方差决定。分析表明,传统正则化引入密集且持久的偏差,改变稳态条件,导致次优策略;而基于协方差的方法仅对高协方差词元进行选择性正则化,并在调节系数退火时实现渐近无偏。这些结果为LLM后训练中的熵控制提供了原则性指导,对扩展强化学习至更大模型和更复杂推理任务具有重要意义。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become a key approach for enhancing reasoning in large language models (LLMs), yet scalable training is often hindered by the rapid collapse of policy entropy, which leads to premature convergence and performance saturation. This paper provides a comparative theoretical analysis of two entropy control strategies: traditional entropy regularization and the recently proposed covariance-based mechanism. We establish a unified framework for entropy dynamics under softmax parameterization, showing that entropy change is governed by the covariance between log-probabilities and logit updates. Our analysis reveals that traditional entropy regularization introduces a dense, persistent bias that modifies the stationary condition, leading to suboptimal policies, while covariance-based methods selectively regularize a sparse subset of high-covariance tokens and achieve asymptotic unbiasedness when the regularization coefficient is annealed. These results provide principled guidelines for entropy control in LLM posttraining, with implications for scaling RL to larger models and more complex reasoning tasks.

强化学习大模型熵控制理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。