针对稀有标签检测难题,提出兼顾标注难度与类别特异性能力的聚合模型。
A Model for Imbalanced Label Aggregation: A Focus on Minority-Class Detection
- 融合项目难度与类别依赖的标注者能力,建模更真实的人工标注行为
- 在33个真实数据集上,稀有类别召回率显著领先,平衡准确率仍具竞争力
- 适合需要精准识别罕见标签的工业质检、医疗诊断等场景
我们研究了以类别依赖标注准确率为焦点的不平衡众包问题,该问题在实际检测系统中至关重要——操作意义最大的标签往往最稀有。在此设定下,标注者可能对两类标签均可靠、均不可靠、仅擅长多数类或仅擅长少数类。现有模型仅部分解决此问题:要么捕捉类别依赖错误却忽略项目难度,要么建模难度却未考虑类别依赖错误。为填补这一空白,我们提出一种生成式聚合模型,同时结合项目难度与类别依赖的标注者能力,允许标注者能力和项目难度在不同类别间变化。我们重新审视了类别不平衡情形下的康多塞陪审团定理,并证明多数投票会渐近保持原始类别比例。我们在33个真实众包数据集上评估该模型,涵盖图像与文本等多类别任务,以及大规模标注与大规模项目两种场景。在各类设置中,该模型始终在稀有类别召回率上表现最优,同时保持良好的平衡准确率,特别适用于以稀有标签恢复为主要目标的任务。
原文摘要 · Abstract (English)
We study imbalanced crowdsourcing with a focus on class-dependent annotator accuracy, a setting that, to the best of our knowledge, remains relatively underexplored despite its importance in real-world inspection systems where the labels of greatest operational importance are also the rarest ones. In this setting, annotators may be reliable on both classes, unreliable on both classes, majority-class specialists, or minority-class specialists. Existing models only partially address this problem: they either capture class-dependent errors but ignore item difficulty, or they model item difficulty without capturing class-dependent errors. To fill this gap for imbalanced datasets in crowdsourcing, we introduce a generative aggregation model combining item difficulty with class-dependent annotator competence. The model allows both annotator abilities and item difficulties to vary across classes. We then revisit Condorcet's Jury Theorem in the class-imbalanced setting. We also show that majority voting asymptotically preserves the underlying class proportion. We evaluate our model on $33$ real-world crowdsourcing datasets, covering multiclass tasks such as images and text, as well as two large-scale regimes: large-scale annotation datasets, with many annotations per item, and large-scale item datasets, with a large number of annotated instances. Across these diverse settings, our model consistently achieves the highest minority recall while remaining competitive in balanced accuracy, making it particularly relevant when rare-label recovery is the primary objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。