arXiv:2605.11538cs.CLcs.AI2026-05ACL

用高斯核动态抑制极端更新,稳定大模型推理训练

Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting

论文配图:Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
图 1 · 摘自论文原文
  • 基于概率与优势的协方差设计权重,自动调节更新强度
  • 在多个推理基准上优于GRPO,且训练中熵更稳定
  • 无需调参,适合追求训练稳定的LLM研究者

组相对策略优化(GRPO)已成为提升大语言模型推理能力的有前景方法。然而,其在训练中难以有效平衡探索与利用,常导致性能不佳。基于熵变化受令牌概率与对应优势协方差支配的理论洞察,我们提出一种无需超参数的协方差加权优化方法,通过高斯核动态下调极端令牌级更新。该方法自动缓解探索-利用权衡带来的不稳定性,同时保留有效学习信号。大量实验证明,本方法在下游推理基准上优于GRPO,且能有效稳定训练过程中的熵。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has emerged as a promising approach for improving the reasoning capabilities of large language models. However, it struggles to effectively balance the tradeoff between exploration and exploitation during training, often resulting in suboptimal performance. Motivated by the theoretical insight that changes in entropy are governed by the covariance between token probabilities and their corresponding advantages, we propose a hyperparameter-free, covariance-weighted optimization method that dynamically down-weights extreme token-level updates via a Gaussian kernel. This approach automatically reduces the instability caused by exploration-exploitation trade-off while preserving informative learning signals. Extensive empirical evaluations show that our approach improves downstream performance across reasoning benchmarks compared with GRPO, and effectively stablizes entropy as training progresses.

强化学习大模型训练策略优化熵控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。