把人工标注看作测量过程,分解出四种变异来源。
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
- 将标注视为测量,拆解出难度、偏见、噪声和关系对齐四类变异
- 在自然语言推理数据集上验证了四种成分均存在且可量化
- 为理解模型学到了什么提供新工具,适合数据质量研究者
监督学习假设标注数据能准确度量模型应学习的概念。但实际中,人工标注因项目模糊、理解差异和错误引入系统性偏差。现有研究常将所有分歧视作噪声,掩盖了关键差异,限制了对模型真实学习内容的理解。本文将标注重构为测量过程,提出一个统计框架,将标注结果分解为可解释的变异来源:实例难度、标注者偏见、情境噪声和关系对齐。该框架扩展经典测量误差模型,兼顾共性与个体化真值观,提供诊断工具以判断任务更符合哪种误差解释。在多标注者自然语言推理数据集上的应用表明,四种理论成分均有实证支持,并证明了方法的有效性。最后讨论其对数据驱动机器学习的启示,指出该方法可推动标注科学的系统化发展。
原文摘要 · Abstract (English)
Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic variation arising from ambiguous items, divergent interpretations, and simple mistakes. Machine learning research commonly treats all disagreement as noise, which obscures these distinctions and limits our understanding of what models actually learn. This paper reframes annotation as a measurement process and introduces a statistical framework for decomposing labeling outcomes into interpretable sources of variation: instance difficulty, annotator bias, situational noise, and relational alignment. The framework extends classical measurement-error models to accommodate both shared and individualized notions of truth, reflecting traditional and human label variation interpretations of error, and provides a diagnostic for assessing which regime better characterizes a given task. Applying the proposed model to a multi-annotator natural language inference dataset, we find empirical evidence for all four theorized components and demonstrate the effectiveness of our approach. We conclude with implications for data-centric machine learning and outline how this approach can guide the development of a more systematic science of labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。