提出新方法提升大模型后训练中的不确定性判断,让优化更稳定可靠。
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization

- 基于几何感知与奖励校准,重构不确定性信号
- 在多个基准上显著改善后训练性能,梯度波动更可控
- 适合追求鲁棒性与可解释性的模型优化研究者
后训练已成为提升大语言模型推理与对齐能力的核心手段,其中无需评判器的模型能从自身生成输出中实现可扩展学习,但缺乏区分有效信号与噪声的理论机制。近期方法利用响应级别熵作为不确定性信号,调节如GRPO等群体优化方法,但其效果不稳定,且对优化过程的影响机制不明确。本文首次提出原则性框架,将不确定性信号理解为调控梯度方差与学习信号质量的机制。通过理论与实证分析,发现现有熵估计器存在两个关键缺陷:各向异性偏差与校准偏差。为此,我们提出几何感知校准策略优化(GCPO),融合语义分歧的几何度量与基于奖励的校准,使不确定性更贴近学习信号强度。实验在多个基准上验证,该方法更准确追踪梯度变化,持续提升后训练表现。结果强调了设计与优化动态对齐的不确定性信号的重要性,为稳健后训练提供原则性视角。
原文摘要 · Abstract (English)
Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from model-generated outputs but lack principled mechanisms to distinguish informative from noisy signals. Recent approaches leverage response-level measures as uncertainty signals to regulate group-based optimization methods such as GRPO. Yet their empirical success remains unstable and unclear in how they influence optimization dynamics. In this paper, we provide, to our knowledge, the first principled formulation that interprets uncertainty signals as mechanisms for characterizing and regulating gradient variance and learning signal quality. Based on both empirical and theoretical analysis, we identify two critical gaps of current entropy-based estimators: The anisotropic gap and The calibration gap. Motivated by this analysis, we propose Geometric-aware Calibrated Policy Optimization (GCPO), a novel framework integrating geometry-aware measures to capture semantic disagreement with reward-based calibration to align uncertainty with learning signal strength. Experiments on multiple benchmarks show that our approach more faithfully tracks gradient variability and consistently improves post-training performance. Our results highlight the importance of designing uncertainty signals that are aligned with optimization dynamics, offering a principled perspective for robust post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。