选对标注一致性度量,让NLP人工标注更可靠。
Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation
- 按任务类型分类选择合适的标注者间一致性指标
- 指出标签不平衡和缺失数据会降低可靠性估计
- 推荐用置信区间和分歧模式分析提升透明度
人工标注仍是自然语言处理中可靠可解释数据的基础。随着标注与评估任务从类别标注扩展到分割、主观判断和连续评分,标注者间一致性的衡量日益复杂。本文梳理了标注者间一致性(IAA)在NLP及相关领域的概念化与应用,阐明常见方法的假设与局限。我们按任务类型组织一致性度量,并讨论标签不平衡与缺失数据如何影响可靠性估计。此外,强调清晰透明报告的最佳实践,包括使用置信区间和分析分歧模式。论文旨在为选择与解读一致性度量提供指南,推动NLP中人工标注与评估的更一致、可复现性。
原文摘要 · Abstract (English)
Human annotation remains the foundation of reliable and interpretable data in Natural Language Processing (NLP). As annotation and evaluation tasks continue to expand, from categorical labelling to segmentation, subjective judgment, and continuous rating, measuring agreement between annotators has become increasingly more complex. This paper outlines how inter-annotator agreement (IAA) has been conceptualised and applied across NLP and related disciplines, describing the assumptions and limitations of common approaches. We organise agreement measures by task type and discuss how factors such as label imbalance and missing data influence reliability estimates. In addition, we highlight best practices for clear and transparent reporting, including the use of confidence intervals and the analysis of disagreement patterns. The paper aims to serve as a guide for selecting and interpreting agreement measures, promoting more consistent and reproducible human annotation and evaluation in NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。