通过解释分析人类在自然语言推理中的标注差异,发现理由一致时标签仍可能不同。
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
- 用解释分类法LiTEx分析标注者推理模式与标签分歧的关系。
- 发现标签不一致但解释相似的案例,说明表面分歧未必反映理解差异。
- 揭示个体标注偏好,提醒别把标签当唯一真实标准。
自然语言推理(NLI)数据集常存在人类标注差异。为深入理解这些差异,本文借助基于解释的方法,以自由文本解释为视角分析标注者决策背后的推理过程。采用LiTEx分类体系,将英文解释划分为推理类别。以往研究多关注标签一致但解释不同的情况,本文拓展至标签与解释双重分歧的分析。我们在两个NLI数据集上应用LiTEx,从三方面对标注差异进行对齐:NLI标签一致性、解释语义相似度及分类类别一致性,并引入标注者选择偏见作为复合因素。结果显示,部分标注者虽在标签上分歧,却提供高度相似的解释,表明表层不一致可能掩盖深层理解的一致性。同时,分析揭示了标注者在解释策略与标签选择上的个体偏好。结果表明,推理类别的一致性比标签一致性更能反映解释的语义相似性。研究强调了基于推理解释的丰富性,也警示不应将标签简单视为绝对真实。
原文摘要 · Abstract (English)
Natural Language Inference (NLI) datasets often exhibit human label variation. To better understand these variations, explanation-based approaches analyze the underlying reasoning behind annotators' decisions. One such approach is the LiTEx taxonomy, which categorizes free-text explanations in English into reasoning categories. However, previous work applying LiTEx has focused on within-label variation: cases where annotators agree on the NLI label but provide different explanations. This paper broadens the scope by examining how annotators may diverge not only in the reasoning category but also in the labeling. We use explanations as a lens to analyze variation in NLI annotations and to examine individual differences in reasoning. We apply LiTEx to two NLI datasets and align annotation variation from multiple aspects: NLI label agreement, explanation similarity, and taxonomy agreement, with an additional compounding factor of annotators' selection bias. We observe instances where annotators disagree on the label but provide similar explanations, suggesting that surface-level disagreement may mask underlying agreement in interpretation. Moreover, our analysis reveals individual preferences in explanation strategies and label choices. These findings highlight that agreement in reasoning categories better reflects the semantic similarity of explanations than label agreement alone. Our findings underscore the richness of reasoning-based explanations and the need for caution in treating labels as ground truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。