arXiv:2605.11775cs.LGcs.CL2026-05

提出熵极性机制,实现对大模型强化学习中探索与利用的精细控制。

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control

论文配图:Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
图 1 · 摘自论文原文
  • 引入熵极性概念,量化每个词元对策略熵的扩张或收缩影响。
  • 实验证明熵极性可准确预测熵变化,正负分支协同提升性能。
  • 新算法PAPO动态调节优化压力,适合需要高效训练的复杂任务。

策略熵已成为理解与控制基于可验证奖励的大语言模型强化学习(RLVR)中探索行为的基础度量。然而,现有熵感知方法主要通过全局目标调控熵,而采样策略更新如何在词元层面重塑策略熵的机制仍不清晰。本文建立了一个关于RLVR中熵力学的理论框架,首次给出熵变化的一阶近似,推导出熵极性——一个能预测采样更新对熵扩张或收缩程度的符号化词元级量。分析揭示结构不对称性:增强高频高概率词元会引发熵收缩趋势,而扩张趋势通常需低概率样本或更强分布修正。实验表明,熵极性可可靠预测熵变化;正负极性分支在保持探索的同时强化利用方面发挥互补作用。基于此,我们提出极性感知策略优化(PAPO),保留双极性分支,并通过优势重加权实现熵控制。以经验熵轨迹作为在线阶段信号,PAPO自适应地在熵扩张与收缩更新间重新分配优化压力。在数学推理与智能体基准上的实验显示,PAPO持续优于主流基线,兼具更高训练效率与显著奖励提升。

原文摘要 · Abstract (English)

Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (RLVR) for LLMs. However, existing entropy-aware methods mainly regulate entropy through global objectives, while the token-level mechanism by which sampled policy updates reshape policy entropy remains underexplored. In this work, we develop a theoretical framework of entropy mechanics in RLVR. Our analysis yields a first-order approximation of the entropy change, giving rise to entropy polarity, a signed token-level quantity that predicts how much a sampled update expands or contracts entropy. This analysis further reveals a structural asymmetry: reinforcing frequent high-probability tokens triggers contraction tendencies, whereas expansive tendencies typically require lower-probability samples or stronger distributional correction. Empirically, we show that entropy polarity reliably predicts entropy changes, and that positive and negative polarity branches play complementary roles in preserving exploration while strengthening exploitation. Building on these insights, we propose Polarity-Aware Policy Optimization (PAPO), which preserves both polarity branches and implements entropy control through advantage reweighting. With the empirical entropy trajectory as an online phase signal, PAPO adaptively reallocates optimization pressure between entropy-expanding and entropy-contracting updates. Experiments on mathematical reasoning and agentic benchmarks show that PAPO consistently outperforms competitive baselines, while delivering superior training efficiency and substantial reward improvements.

强化学习大模型熵控制策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。