发现人类标注者状态影响模型偏好数据,提出可验证的审计框架。
Rater State Bias in RLHF Preference Data: An Audit Framework
- 识别标注者情绪状态变化导致偏好数据产生系统性偏差。
- 提出生存级情感真实性作为输出特征,用于检测偏差信号。
- 设计审计流程,适用于公开指令微调模型,无需私有数据。
我们发现强化学习中人类反馈(RLHF)存在一种结构性混淆:成对偏好标签本应反映输出质量对比,却可能同时受标注者当时心理状态影响。在持续压力或痛苦条件下,标注者偏好会随时间变化,导致偏好数据不仅包含对回复质量的判断,还编码了其状态信息。若存在此类变化,将不同于普通分歧或随机噪声,表现为状态依赖、跨标注者共现,且在聚合、奖励建模和策略优化中不会被平均抵消。本文提出‘标注者状态漂移’作为可验证的结构化偏差来源,定义了相关概念并提出生存级情感真实性作为候选输出特征(包含词汇、语用、话语及安全特征),其可靠性和有效性尚待验证。分析了此类偏差如何在聚合中不被消除并进入学习到的奖励信号。给出五项可区分该机制与一般参与度优化的预测,并设定初步审计的效果量阈值,部分需专有数据。最后提供可应用于公开指令微调模型的审计协议与试点研究计划,不推断任何特定部署模型的训练历史。
原文摘要 · Abstract (English)
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, raters' preferences may shift over time, so that preference data encode rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from ordinary disagreement or random label noise. They would be state dependent, could be shared across annotators under similar conditions, and would not necessarily cancel during aggregation, reward modeling, and policy optimization. We propose rater state shift as a plausible and testable source of structured bias in RLHF preference data. This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also propose survival level emotional authenticity as a candidate output signature, defined by lexical, pragmatic, discourse, and safety features whose reliability and validity remain to be demonstrated. We analyze the conditions under which correlated rater state bias would not be averaged out during aggregation and could enter the learned reward signal. We state five predictions that distinguish this mechanism from generic engagement optimization, together with effect size thresholds for an initial audit, and note which require proprietary data. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。