arXiv:2601.09065cs.CL2026-01被引 8

将标注分歧视为观点差异,而非噪声,提升NLP任务的可信度。

Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP

  • 构建分歧来源的通用分类体系,涵盖数据、任务与标注者因素。
  • 提出统一框架,支持显式建模标注者间关系与分歧结构。
  • 适合关注主观任务公平性与可解释性的研究人员参考。

在自然语言处理中,标注分歧广泛存在于主观与模糊任务(如毒性检测、立场分析)中。早期方法将分歧视为需剔除的噪声,而近期研究则将其视为反映不同解读与视角的有意义信号。本文提供一种统一视角,系统梳理分歧感知的NLP方法。首先建立跨领域的分歧来源分类体系,涵盖数据、任务与标注者三类因素;其次基于预测目标与聚合结构,归纳建模方法,揭示从共识学习向显式建模分歧及标注者间结构关系的演进趋势;进一步综述预测性能与标注行为的评估指标,指出当前公平性评估多为描述性而非规范性;最后提出开放挑战,包括整合多种变异源、发展分歧感知的可解释性框架,以及权衡视角主义建模的实际代价。

原文摘要 · Abstract (English)

Annotator disagreement is widespread in NLP, particularly for subjective and ambiguous tasks such as toxicity detection and stance analysis. While early approaches treated disagreement as noise to be removed, recent work increasingly models it as a meaningful signal reflecting variation in interpretation and perspective. This survey provides a unified view of disagreement-aware NLP methods. We first present a domain-agnostic taxonomy of the sources of disagreement spanning data, task, and annotator factors. We then synthesize modeling approaches using a common framework defined by prediction targets and pooling structure, highlighting a shift from consensus learning toward explicitly modeling disagreement, and toward capturing structured relationships among annotators. We review evaluation metrics for both predictive performance and annotator behavior, and noting that most fairness evaluations remain descriptive rather than normative. We conclude by identifying open challenges and future directions, including integrating multiple sources of variation, developing disagreement-aware interpretability frameworks, and grappling with the practical tradeoffs of perspectivist modeling.

标注分歧主观任务可解释性评价方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。