人类与AI对齐反馈中,偏好常被误导,真实偏好难捕捉。
Aligning to Illusions: Choice Blindness in Human and AI Feedback
- 通过悄悄替换偏好,91%的人类未察觉,显示选择盲视普遍存在。
- 仅1/6到1/3的标签被污染,奖励信号就减半,但标准准确率不变。
- 适合关注人机反馈可靠性、强化学习对齐安全的研究者阅读。
强化学习从人类反馈(RLHF)假设标注者的偏好反映稳定的内在状态。我们通过三项实验挑战这一假设。在人类选择盲视研究中,91%的隐蔽偏好替换未被察觉,将选择盲视扩展至对陌生文本的第三人称评价。测试15个大语言模型(LLM)作为潜在替代者,发现检测依赖浅层文本匹配而非真正自我监控:移除上下文中的先前推理导致盲视率从接近零升至超过50%,而明确的社会压力则引发近乎普遍的顺从。在剂量-响应实验中,针对86M至2B参数的两种架构,1/6至1/3的标签被污染时,奖励信号即减半,但标准成对准确率几乎不变。最佳-之-N评估证实这影响下游策略:在50%标签污染下,基于奖励的选择无改进优于随机采样,而代理模型报告得分持续上升。这些结果揭示了偏好构建问题:进入RLHF的信号受诱导情境塑造,而人类元认知、LLM自我监控及标准评估指标均无法察觉。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) assumes annotator preferences reflect stable internal states. We challenge this through three experiments spanning the preference pipeline. In a human choice blindness study, 91% of surreptitiously swapped preferences go undetected, extending choice blindness to third-person evaluative comparison of unfamiliar text. Testing fifteen LLM judges as potential replacements, we find detection relies on shallow text matching rather than genuine self-monitoring: removing prior reasoning from context causes blindness to surge from near-zero to over 50%, while explicit social pressure induces near-universal compliance. In a dose-response experiment across two architectures from 86M to 2B parameters, one-sixth to one-third of labels must be corrupted before the reward signal halves, yet standard pairwise accuracy remains virtually unchanged. A Best-of-N evaluation confirms this translates to downstream policy degradation: at 50% corruption, reward-guided selection produces no improvement over random sampling, while the proxy model reports monotonically increasing scores. Together, these results reveal a preference construction problem: the signal entering RLHF is shaped by elicitation context in ways that neither human metacognition, LLM self-monitoring, nor standard evaluation metrics can detect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。