奖励欺骗让大模型误判自身信心,影响可靠性。
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs

- 用诱导错误认同的奖励机制微调模型,模拟奖励作弊
- 模型不确定性量化能力下降,误差率上升0.6%至1.0%
- 即使校准后仍残留偏差,适合关注模型可信度的研究者
现代大语言模型越来越多地通过人类反馈强化学习(RLHF)或相关奖励优化方法进行微调。尽管这类方法提升了模型的感知帮助性,但我们研究了迎合性奖励信号是否损害了校准性——这对可靠不确定性量化至关重要。我们在Qwen3-8B上进行了三种微调策略:无微调(基础模型)、在TriviaQA上的中性监督微调(SFT),以及诱导迎合性的组相对策略优化(GRPO),后者奖励模型对预设错误答案的认同。在涵盖五个学科领域的1,000个MMLU题目上,使用置信区间和置换检验评估,发现迎合性GRPO导致系统性校准恶化:相对于基础模型,ECE上升0.006;相对于中性SFT,MCE上升0.010,但该效应在当前训练预算下未达统计显著性(p=0.41)。对三类模型施加事后矩阵校准后,ECE降低40%–64%,准确率提升1.5–3.0个百分点。然而,迎合性模型的校准后ECE仍高于中性对照组(0.042 vs. 0.037),表明奖励引发的校准偏差会留下结构性残余。本研究建立了评估奖励作弊对校准影响的方法,并推动校准感知的训练目标设计。
原文摘要 · Abstract (English)
Modern large language models (LLMs) are increasingly fine-tuned via reinforcement learning from human feedback (RLHF) or related reward optimisation schemes. While such procedures improve perceived helpfulness, we investigate whether sycophantic reward signals degrade calibration -- a property essential for reliable uncertainty quantification. We fine-tune Qwen3-8B under three regimes: no fine-tuning (base), neutral supervised fine-tuning (SFT) on TriviaQA, and sycophancy-inducing Group Relative Policy Optimisation (GRPO) that rewards agreement with planted wrong answers. Evaluating on $1{,}000$ MMLU items across five subject domains with bootstrap confidence intervals and permutation testing, we find that \textbf{sycophantic GRPO produces consistent directional calibration degradation} -- ECE rises by $+0.006$ relative to the base model and MCE increases by $+0.010$ relative to neutral SFT -- though the effect does not reach statistical significance ($p = 0.41$) at this training budget. Post-hoc matrix scaling applied to all three models reduces ECE by $40$--$64\%$ and improves accuracy by $1.5$--$3.0$ percentage points. However, the sycophantic model retains the highest post-scaling ECE relative to the neutral SFT control ($0.042$ vs.\ $0.037$), suggesting that reward-induced miscalibration leaves a structured residual even after affine correction. These findings establish a methodology for evaluating the calibration impact of reward hacking and motivate calibration-aware training objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。