arXiv:2510.12817cs.CLcs.AI2025-10ACL被引 4

保留人类标注差异,让模型更贴近真实多元的人类价值观。

From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP

  • 将人类标注差异视为信号而非噪声,体现多元视角
  • 现有偏好数据集常合并多标注,导致观点单一化
  • 建议在数据构建中保留标注多样性,用于模型对齐与安全评估

人类标注差异(HLV)指标注中合法的分歧,反映人类视角的多样性而非单纯错误。长期以来在NLP中被视为需消除的噪声,近年被重新视为提升模型鲁棒性的信号。随着大语言模型和基于人类反馈的后训练方法兴起,HLV的重要性日益凸显。然而,现有偏好学习数据集通常将多个标注合并为单一标签,人为制造共识,抹平了多元观点。保留HLV不仅有助于多元对齐,也对社会技术安全性评估至关重要,需结合人类互动与社会背景评估模型行为。本文主张将保留HLV作为内在价值(Selbstzweck),分析现有数据集的局限性,并提出在数据构建中融入HLV的具体策略,以更好保存多元人类价值。

原文摘要 · Abstract (English)

Human Label Variation (HLV) refers to legitimate disagreement in annotation that reflects the diversity of human perspectives rather than mere error. Long treated in NLP as noise to be eliminated, HLV has only recently been reframed as a signal for improving model robustness. With the rise of large language models (LLMs) and post-training methods such as human feedback-based alignment, the role of HLV has become increasingly consequential. Yet current preference-learning datasets routinely collapse multiple annotations into a single label, flattening diverse perspectives into artificial consensus. Preserving HLV is necessary not only for pluralistic alignment but also for sociotechnical safety evaluation, where model behavior must be assessed in relation to human interaction and societal context. This position paper argues that preserving HLV as an embodiment of human pluralism must be treated as a Selbstzweck, an intrinsic value in itself. We analyze the limitations of existing preference datasets and propose actionable strategies for incorporating HLV into dataset construction to better preserve pluralistic human values.

人类标注模型对齐多元价值后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。