人类偏好反馈受注意力限制,标准模型会误判真实偏好。
Attention Limited Reward Learning
- 用有限注意力机制建模人类比较行为,区分真实价值接近与感知困难。
- 被动数据无法区分奖励、注意力和默认倾向,易导致错误排名。
- 建议将反馈视为注意力受限的测量信号,适合对对齐机制敏感的研究者。
成对人类比较是现代AI系统学习人类偏好的主要方式。当前的强化学习人类反馈(RLHF)等对齐流程通常使用Bradley-Terry log-odds模型,其中选择概率由潜在奖励差异决定。本文基于理性不注意理论,提出一个简化模型,认为每个标签由低容量评估通道生成。该模型分离了标准奖励建模中混淆的两种模糊性:候选项可能因真实价值相近而难分胜负,也可能因注意能力有限导致特征难以察觉。我们发现,注意力限制会根本性扭曲成对比较所揭示的信息。特别是,被动比较数据通常无法区分奖励、注意力和默认倾向;异质注意力可能导致标准Bradley-Terry模型恢复出误导性排序。分析表明,学习效果取决于每条标签携带的注意信息量,而非标签数量本身。在Chatbot Arena上对语言模型对的人类投票案例研究中,发现了预测中的循环成分,其幅度超过采样噪声,任何标量奖励都无法表示;另一组感知比较研究表明,反应时间和注视轨迹包含了标签未反映的差距信息。这一视角提示:人类反馈不应视为直接的偏好揭示,而应看作注意力受限的测量过程——弱偏好信号可能反映隐藏的评估难度,而非真正的无差别。
原文摘要 · Abstract (English)
Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipelines typically model such comparisons with Bradley--Terry log-odds, where choice probabilities are governed by latent reward differences. This paper examines what this assumption misses through a reduced-form model motivated by rational inattention, in which each label is generated by a low-capacity evaluation channel. The model separates two forms of ambiguity that standard reward modeling tends to conflate: a comparison may be difficult because the two candidates are genuinely close in value, or because the relevant distinction is hard to detect under limited attention. We show that limited attention can fundamentally distort what pairwise comparisons reveal. In particular, passive comparison data cannot generally distinguish reward, attention, and default tendencies, and heterogeneous attention can make standard Bradley--Terry reward modeling recover misleading rankings. Our analysis shows that learning is governed not by the raw number of labels, but by the amount of attended information each label carries. A case study on human votes over language-model pairs from Chatbot Arena exhibits the predicted signature, a cyclic component of the comparison data that exceeds sampling noise and that no scalar reward can represent; a second case study on perceptual comparisons shows that response times and gaze carry gap information that the labels do not. This perspective suggests that human feedback should be treated not as direct revealed preference, but as an attention-limited measurement process: a weak preference signal may reflect hidden evaluation difficulty rather than genuine indifference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。