arXiv:2602.22723cs.CL2026-02

研究人类标注差异对隐式语篇关系识别的影响

Human Label Variation in Implicit Discourse Relation Recognition

  • 比较分布预测与个体标注者建模两种方法
  • 分布预测模型在模糊任务中表现更稳定
  • 认知复杂度高的案例导致标注不一致

许多自然语言处理任务缺乏单一标准答案,因人类判断反映不同视角。为捕捉这种差异,已有模型尝试预测完整标注分布而非多数标签;而视角主义模型则旨在复现单个标注者的理解。本文在高度模糊的隐式语篇关系识别(IDRR)任务上对比这些方法。实验表明,除非降低模糊性,否则现有标注者特定模型表现不佳;而基于标签分布训练的模型能产生更稳定的预测。进一步分析显示,频繁出现的认知复杂案例是导致人类判断不一致的主要原因,这对视角主义建模在IDRR中的应用构成挑战。

原文摘要 · Abstract (English)

There is growing recognition that many NLP tasks lack a single ground truth, as human judgments reflect diverse perspectives. To capture this variation, models have been developed to predict full annotation distributions rather than majority labels, while perspectivist models aim to reproduce the interpretations of individual annotators. In this work, we compare these approaches on Implicit Discourse Relation Recognition (IDRR), a highly ambiguous task where disagreement often arises from cognitive complexity rather than ideological bias. Our experiments show that existing annotator-specific models perform poorly in IDRR unless ambiguity is reduced, whereas models trained on label distributions yield more stable predictions. Further analysis indicates that frequent cognitively demanding cases drive inconsistency in human interpretation, posing challenges for perspectivist modeling in IDRR.

语篇分析标注差异模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。