arXiv:2605.27996cs.AI2026-05被引 1

修复奖励模型偏见时,可能把压力转移到其他相关偏见上,导致新问题。

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

论文配图:Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
图 1 · 摘自论文原文
  • 通过分析评估与训练分布差异,揭示偏见转移的机制
  • 实证发现长度惩罚引发模型过度自信,准确率下降
  • 建议结合政策诱导分布评估多维度偏见,避免误判

单一维度的奖励模型偏见缓解(如减少对长度、迎合性或风格的依赖)可能导致优化压力转向相关代理变量,这种现象称为奖励偏见替代。该失败源于评估分布与策略生成分布之间的测量-优化差距。我们构建缓解结果的分类体系,证明在任何审计评分下(包括排名准确率和胜率),成功缓解、偏见替代和过度修正产生相同观测结果,即使拥有真实奖励的预言机也难以区分。我们调查的已有偏好学习缓解方法均未提供认证成功缓解的证据。通过在评估中引入策略诱导分布并跟踪多个偏见,可真正关闭这一差距,并提出可操作的缓解方法与基准建议。我们在语言模型RLHF中验证了偏见替代:在GRPO训练中加入长度惩罚虽压缩了响应,却将优化压力转移到置信度校准,导致模型过度自信,自由形式事实准确性下降。此外,一种已发表的长度去偏算子虽在审计分布上消除长度-奖励相关性,但在最佳选择-N策略下,仍使三个出错率最高的奖励模型重新引入偏见;长度与迎合性的耦合方向在人类与大模型评委不一致时甚至反转。

原文摘要 · Abstract (English)

Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than eliminate it, a failure mode we call reward bias substitution. The failure is enabled by a measurement-versus-optimization gap between audit and policy-induced distributions during mitigation evaluation and policy training. We formalize mitigation outcomes into a regime taxonomy and prove that successful mitigation, bias substitution, and overcorrection produce identical observables under any audit-distribution scoring, including ranking accuracy and win-rate, even when granted oracle access to the true reward. Across published preference-learning mitigation work, no method we survey reports the evidence needed to certify successful mitigation. Augmenting evaluation with policy-induced distributions while tracking multiple biases provably closes the gap, and we translate this into actionable prescriptions for mitigation methods and benchmarks. We demonstrate bias substitution in language model RLHF, where a length penalty during GRPO training compresses responses as intended yet redirects optimization pressure onto confidence calibration, driving the policy into overconfidence while factual free-form accuracy falls. We also show a published length-debiasing operator that zeroes reward-length correlation on the audit distribution but reintroduces bias under best-of-N selection on three of four SOTA reward models, and a length-sycophancy coupling whose direction reverses under human-LLM judge disagreement.

RLHF偏见缓解奖励模型优化压力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。