揭示了GSPO中重要性权重与困惑度、熵的内在等价关系。
Rethinking GSPO: The Perplexity-Entropy Equivalence
- 用信息论重新解释GSPO的序列权重,将其等价为困惑度比和熵变指数。
- 实验证明该等价关系在数学推理任务中成立,且能解释训练稳定性。
- 适合关注强化学习权重设计与模型稳定性的研究者阅读。
我们通过建立长度归一化重要性权重与信息论量之间的联系,为GSPO提供了新视角。发现GSPO的序列级权重 $s(θ) = (π_θ/π_{θ_{ ext{old}}})^{1/|y|}$ 可等价表示为逆困惑度比 $ ext{PPL}_{θ_{ ext{old}}}/ ext{PPL}_θ$ 与熵变化的指数 $ ext{exp}(ΔH)$。尽管困惑度-熵关系源于标准定义,但这一观察为理解GSPO提供了有用视角:算法通过困惑度比对策略梯度更新加权,赋予重要性权重以信息论解释。该视角有助于说明GSPO的实验性质,包括通过对数域的几何平均实现方差降低,以及在混合专家模型训练中的稳定性。我们在数学推理任务上通过受控实验验证了数学等价性和方差预测。
原文摘要 · Abstract (English)
We provide a new perspective on GSPO's length-normalized importance ratios by establishing their connection to information-theoretic quantities. We show that GSPO's sequence-level weight $s(θ) = (π_θ/π_{θ_{\text{old}}})^{1/|y|}$ can be equivalently expressed as the inverse perplexity ratio $\text{PPL}_{θ_{\text{old}}}/\text{PPL}_θ$ and as the exponential cross-entropy change $\exp(ΔH)$. While the perplexity-entropy relationship follows from standard definitions, this observation provides a useful lens for understanding GSPO: the algorithm weights policy gradient updates by perplexity ratios, offering an information-theoretic interpretation of the importance weights. This perspective helps explain GSPO's empirical properties, including log-domain variance reduction through geometric averaging and stability in training mixture-of-experts models. We validate the mathematical equivalences and variance predictions through controlled experiments on mathematical reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。