arXiv:2606.10569cs.CLcs.AI2026-06

提出人类反馈中的多元共识测量问题,避免因简化判断而丢失文化多样性。

Hidden Consensus:Preference-Validity Compression in Human Feedback

论文配图:Hidden Consensus:Preference-Validity Compression in Human Feedback
图 1 · 摘自论文原文
  • 用马来西亚多视角数据揭示反馈聚合中多元有效回答被压缩为单一奖励的问题
  • 321个偏好事件中79%的提示存在多个主流可接受回复,单胜者机制会丢弃这些有效选项
  • 适合关注AI对齐中文化差异、社会多样性与公平性的研究者和实践者

标准的强化学习人类反馈(RLHF)流程常将异质的人类判断简化为单一标量奖励目标。我们指出,这种简化在结构多元的社会中可能误判对齐程度,因为分歧可能源于文化、历史、语言、地域或规范性基础,而非标注噪声。我们称之为‘偏好-有效性压缩’,即多个合理有效的回应被压缩为单一优化目标。以马来西亚为诊断场景,我们通过链接提示、回复与可接受性判断的偏好事件,分析了类似RLHF的反馈聚合机制。在20名参与者和107个三重标注提示的321个偏好事件中,79%的提示包含多个主流支持的回应,而单胜者聚合会将其丢弃;当所有主流支持选项被纳入考量时,顶级回应间的明显优势差距显著缩小。参与者频繁选择多个可接受回复,被丢弃的回应也明显反映一致的本地化、实用性或文化框架。结果表明,当前多数聚合方法衡量的是‘最大可接受性’而非‘多元对齐’。我们将此视为测量有效性问题,主张未来对齐方法应满足‘有效性保持一致性’,在多元有效解释框架间保持稳定,而非将其压缩为单一奖励目标。

原文摘要 · Abstract (English)

Standard RLHF pipelines often reduce heterogeneous human judgments into a single scalar reward target. We argue that this reduction can mis-measure alignment in structurally plural societies, where disagreement may reflect culturally, historically, linguistically, regionally, or normatively grounded interpretations rather than annotation noise. We call this failure Preference-Validity Compression, the collapse of multiple plural-valid response options into a single optimization target. Using Malaysia as a diagnostic setting, we analyze RLHF-style feedback aggregation through preference events linking prompts, responses, and acceptability judgments across interpretive frames. Across 321 preference events from 20 participants and 107 trio-annotated prompts, 79% of prompts contain more than one majority-supported response that single-winner aggregation would discard, and apparent dominance gaps between top responses diminish when all majority-supported options are considered. Participants frequently select multiple acceptable responses, and discarded responses demonstrably reflect coherent local, practical, or cultural frames. These findings show that majority aggregation in this corpus measures argmax acceptability rather than plural alignment. We treat this as a measurement-validity issue and argue that future alignment methods should satisfy Validity-Preserving Consistency, remaining stable across plural-valid interpretive frames rather than collapsing them into a single reward target.

人类反馈对齐偏差文化多样性评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。