arXiv:2602.03392cs.LGcs.AI2026-02被引 7

揭示大模型强化微调中熵的变化规律,助力平衡探索与利用。

On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models

  • 建立熵变化的理论框架,推导单次逻辑值更新下的熵变公式。
  • 提出熵判别裁剪方法,实验证明能有效优化探索-利用平衡。
  • 为现有熵调控方法提供统一解释视角,适合关注训练稳定性的研究者。

熵是衡量大语言模型输出多样性的关键指标,有助于理解其探索能力。尽管近期研究越来越关注在强化微调(RFT)中监控和调整熵以平衡探索与利用,但对熵动态变化的系统性理论理解仍不充分。本文建立了分析RFT过程中熵动态变化的理论框架,从单次逻辑值更新下的熵变判别表达式出发,推导出一阶熵变表达式,并可进一步扩展至分组相对策略优化(GRPO)的更新公式。由此得出的推论与洞见启发了熵控制方法的设计,也为现有研究中的各类基于熵的方法提供了统一解释视角。我们通过实证验证了分析的核心结论,并展示了所提出的熵判别裁剪方法的有效性。本研究为理解大模型微调的训练动态提供了新见解,为优化探索-利用平衡提供了理论支持与实用策略。

原文摘要 · Abstract (English)

Entropy serves as a critical metric for measuring the diversity of outputs generated by large language models (LLMs), providing valuable insights into their exploration capabilities. While recent studies increasingly focus on monitoring and adjusting entropy to better balance exploration and exploitation in reinforcement fine-tuning (RFT), a principled understanding of entropy dynamics during this process is yet to be thoroughly investigated. In this paper, we establish a theoretical framework for analyzing the entropy dynamics during the RFT process, which begins with a discriminant expression that quantifies entropy change under a single logit update. This foundation enables the derivation of a first-order expression for entropy change, which can be further extended to the update formula of Group Relative Policy Optimization (GRPO). The corollaries and insights drawn from the theoretical analysis inspire the design of entropy control methods, and also offer a unified lens for interpreting various entropy-based methods in existing studies. We provide empirical evidence to support the main conclusions of our analysis and demonstrate the effectiveness of the derived entropy-discriminator clipping methods. This study yields novel insights into RFT training dynamics, providing theoretical support and practical strategies for optimizing the exploration-exploitation balance during LLM fine-tuning.

强化微调熵分析模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。