通过分析标注者身份,更准确预测人类对主观内容的分歧。
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
- 用可学习的权重动态评估不同人口特征对分歧的影响
- 在两个基准上达成0.75的分歧追踪相关性,显著优于大模型方法
- 适合关注公平性与多样性建模的研究者与实践者
人类标注主观内容时会产生分歧,这种分歧并非噪声,而是由标注者的社会身份和生活经历所塑造的真实视角差异。当前主流做法仍将其简化为单一多数标签,近期基于大语言模型的方法也未能改善这一问题:我们证明,即使采用思维链推理,提示式大模型仍无法恢复人类分歧的结构。为此提出DiADEM,一种神经架构,能学习“每个人口维度在预测谁会分歧及分歧内容中有多重要”。DiADEM通过每个人口维度的投影编码标注者,由可学习的重要性向量α控制,利用互补拼接与哈达玛积融合标注者与内容表征,并采用新型项级分歧损失直接惩罚预测的标注方差。在DICES对话安全性和VOICED政治冒犯性两个基准上,DiADEM在标准与视角导向指标上均显著优于LLM作为评判者及神经基线模型,实现0.75的分歧追踪相关性(r=0.75)。学习到的α权重显示,种族与年龄在两个数据集中始终是最具影响力的分歧驱动因素。结果表明,要真实反映人类解释多样性,必须显式建模标注者的身份而非仅关注其标注内容。
原文摘要 · Abstract (English)
When humans label subjective content, they disagree, and that disagreement is not noise. It reflects genuine differences in perspective shaped by annotators' social identities and lived experiences. Yet standard practice still flattens these judgments into a single majority label, and recent LLM-based approaches fare no better: we show that prompted large language models, even with chain-of-thought reasoning, fail to recover the structure of human disagreement. We introduce DiADEM, a neural architecture that learns "how much each demographic axis matters" for predicting who will disagree and on what. DiADEM encodes annotators through per-demographic projections governed by a learned importance vector $\boldsymbolα$, fuses annotator and item representations via complementary concatenation and Hadamard interactions, and is trained with a novel item-level disagreement loss that directly penalizes mispredicted annotation variance. On the DICES conversational-safety and VOICED political-offense benchmarks, DiADEM substantially outperforms both the LLM-as-a-judge and neural model baselines across standard and perspectivist metrics, achieving strong disagreement tracking ($r{=}0.75$ on DICES). The learned $\boldsymbolα$ weights reveal that race and age consistently emerge as the most influential demographic factors driving annotator disagreement across both datasets. Our results demonstrate that explicitly modeling who annotators are not just what they label is essential for NLP systems that aim to faithfully represent human interpretive diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。