区分标注者分歧中的真实信号与噪声,提升数据质量
NUTMEG: Separating Signal From Noise in Annotator Disagreement
- 基于贝叶斯框架融合标注者背景信息,区分系统性分歧与错误
- 在模拟数据中恢复真实标签的准确率优于传统聚合方法
- 适合处理标注者能力差异大、存在合理分歧的任务
NLP模型常依赖人工标注数据进行训练与评估。许多方法通过大量背景、能力、动机各异的标注者众包获取数据,导致标注冲突。传统聚合方法将分歧视为错误并消除,但近期研究指出,对许多任务而言,分歧具有真实意义,应作为信号而非噪声。然而,现有模型极少能有效分离信号与噪声。本文提出NUTMEG,一种新的贝叶斯模型,利用标注者背景信息,剔除噪声标注的同时保留系统性分歧。通过合成数据验证,NUTMEG在存在系统性分歧时比传统聚合方法更有效地恢复真实标签。我们进一步分析了子群体规模、分歧率和垃圾标注率对模型性能的影响。最终实验表明,使用NUTMEG聚合的数据训练下游模型,其性能显著优于传统方法。结果强调了在人工标注数据训练中同时考虑标注者能力与系统性分歧的重要性。
原文摘要 · Abstract (English)
NLP models often rely on human-labeled data for training and evaluation. Many approaches crowdsource this data from a large number of annotators with varying skills, backgrounds, and motivations, resulting in conflicting annotations. These conflicts have traditionally been resolved by aggregation methods that assume disagreements are errors. Recent work has argued that for many tasks annotators may have genuine disagreements and that variation should be treated as signal rather than noise. However, few models separate signal and noise in annotator disagreement. In this work, we introduce NUTMEG, a new Bayesian model that incorporates information about annotator backgrounds to remove noisy annotations from human-labeled training data while preserving systematic disagreements. Using synthetic data, we show that NUTMEG is more effective at recovering ground-truth from annotations with systematic disagreement than traditional aggregation methods. We provide further analysis characterizing how differences in subpopulation sizes, rates of disagreement, and rates of spam affect the performance of our model. Finally, we demonstrate that downstream models trained on NUTMEG-aggregated data significantly outperform models trained on data from traditionally aggregation methods. Our results highlight the importance of accounting for both annotator competence and systematic disagreements when training on human-labeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。