研究标注者分歧根源,揭示模型评估与训练中的隐性偏差。
Diverging Preferences: When do Annotators Disagree and do Models Know?
- 构建十类分歧来源分类体系,发现多数分歧源于任务模糊或表达风格差异。
- 实验证明标准奖励建模与大模型评判方法忽略标注者分歧,导致评估失真。
- 提出识别并缓解分歧影响的方法,助力更公平的模型对齐与评估。
我们研究人类标注偏好数据集中存在的分歧现象。构建涵盖四大类十个子类的分歧来源分类体系,发现多数分歧由任务描述不明确或回答风格差异导致。这一发现挑战了主流奖励建模中将标注分歧视为简单噪声的假设。实验表明,标准奖励建模(如Bradley-Terry)和基于大模型作为评判者的评估方法均未能有效处理标注者间的分歧。这些结果凸显出大模型评估受响应风格等分裂特征显著影响的问题,也对实现多元对齐的大模型提出了挑战。为此,我们提出识别分歧偏好的方法,以减轻其在模型评估和训练中的负面影响。
原文摘要 · Abstract (English)
We examine diverging preferences in human-labeled preference datasets. We develop a taxonomy of disagreement sources spanning ten categories across four high-level classes and find that the majority of disagreements are due to factors such as task underspecification or response style. Our findings challenge a standard assumption in reward modeling methods that annotator disagreements can be attributed to simple noise. We then explore how these findings impact two areas of LLM development: reward modeling training and evaluation. In our experiments, we demonstrate how standard reward modeling (e.g., Bradley-Terry) and LLM-as-Judge evaluation methods fail to account for divergence between annotators. These findings highlight challenges in LLM evaluations, which are greatly influenced by divisive features like response style, and in developing pluralistically aligned LLMs. To address these issues, we develop methods for identifying diverging preferences to mitigate their influence in evaluations and during LLM training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。