发现语言模型奖励机制存在持续偏见,提出可扩展的修正方法。
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
- 通过系统测试五种高质量奖励模型,识别出长度、奉承、过度自信等旧偏见
- 发现新偏见:模型风格偏好和回答顺序依赖,部分偏见可通过简单干预消除
- 提出机制化奖励重塑法,低数据量下有效修复偏见且泛化性强
奖励模型(RMs)在将语言模型(LMs)与人类偏好在线对齐中起关键作用。然而,基于奖励模型的偏好微调易受奖励欺骗影响,导致语言模型学习到不良行为。通过对五种高质量奖励模型(包括最先进模型)进行系统性测量,我们发现尽管已有研究努力,长度偏差、奉承倾向和过度自信等问题依然存在。我们还发现了与模型特定“风格”及回答顺序相关的新型偏见。将奖励模型失败分为可线性干预或难以干预两类,并提出一种简单的后处理干预方法,以缓解由虚假相关引起的低复杂度偏见。所提出的机制化奖励塑造方法在不降低奖励质量的前提下,有效减少目标偏见,且仅需少量标注数据。该方法可扩展至新偏见、模型内部机制,并具有分布外泛化能力。
原文摘要 · Abstract (English)
Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, including the state-of-the-art, we find that issues persist despite prior work with respect to length, sycophancy, and overconfidence. We also discover new issues related to bias toward model-specific ``styles'' and answer-order. We categorize RM failures as tractable or resistant to linear intervention and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. Our proposed mechanistic reward shaping reduces targeted biases without degrading reward quality and while using minimal labeled data. The method is extensible to new biases, model-internal, and generalizes out-of-distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。