解决少量标注者参与任务导致的标注可信度估计难题
AHEAD: Advancing Multi-Class Label Aggregation with Interpretable Cross-Annotator Modeling

- 用图神经网络学习标注者上下文,生成可解释的标注特征
- 在10个真实数据集上平均准确率提升至73.23%,最高增益14.9%
- 适合处理标注者覆盖不全的多类别标注聚合场景
众包标注为自然语言处理、计算机视觉和视频等领域提供了宝贵标注数据。标签聚合旨在从噪声和偏差标注中推断出真实标签,核心在于标注者可靠性估计。然而现有方法面临一个现实瓶颈:多数标注者仅参与少数任务,导致标注者可靠性估计极难实现。本文聚焦更具挑战性的多类别标签聚合,提出AHEAD(跨标注者学习与高置信度标注者引导的标签聚合)框架,通过利用群体级数据推进标注者可靠性估计。具体地,AHEAD首先通过图神经网络学习高维跨标注者上下文,结合个体标注者特征与上下文信息生成多视角互补的标注者嵌入;再将这些嵌入解码为可解释的标注者特定混淆矩阵以拟合观测标签。我们构建包含高置信度标注者的复合目标函数,缓解先前模型面临的无监督训练问题。在涵盖NLP、CV、Video和Audio的10个真实数据集上实验表明,AHEAD显著提升标签准确率,平均准确率从68.75%提升至73.23%,最佳情况提升达14.9%。此外,在最大数据集上的可扩展性实验进一步证明了该方法的整体优势。
原文摘要 · Abstract (English)
Crowdsourced labeling provides valuable labeled data for domains across natural language processing, computer vision, and video. Label aggregation aims to infer latent true labels from noisy and biased annotations, with the key lying in annotator reliability estimation. Despite promising progress, existing approaches struggle with one real-world bottleneck: most individual annotators label only a small subset of tasks, making accurate annotator estimation highly intractable. In this paper, we focus on the considerably more challenging multi-class label aggregation and propose AHEAD (cross-Annotator learning and High-confidEnce Annotator-guideD label aggregation), a cross-annotator learning framework that advances annotator reliability estimation by leveraging the population-level data. Specifically, AHEAD first learns high-dimensional cross-annotator contexts via a graph neural network, deriving multi-view, complementary annotator embeddings by aggregating individual-level annotator features with contextual information. These embeddings are then decoded into interpretable annotator-specific confusion matrices to fit the observed labels. We formulate a composite objective incorporating high-confidence annotators to alleviate the unsupervised training issues faced by prior models. Experiments on 10 real-world datasets spanning NLP, CV, Video, and Audio show that AHEAD substantially improves label accuracy, increasing average accuracy from 68.75% to 73.23%, with gains of up to 14.9% in the best case. Meanwhile, scalability experiments on the largest dataset further demonstrate the overall superiority of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。