arXiv:2512.12870cs.LGcs.AI2025-12

针对标注错误问题,优化选样与标注者分配以提升模型鲁棒性。

Optimal Labeler Assignment and Sampling for Active Learning in the Presence of Imperfect Labels

  • 按标签者能力分配样本,降低每轮最大噪声水平。
  • 新采样策略有效减少标签噪声对模型的影响。
  • 适合标注质量不一的现实场景,如众包标注任务。

主动学习(AL)在标注成本高的应用中备受关注,其通过查询有信息量的样本由标注者(oracle)进行标注。然而,由于标注者准确率参差不齐,标签常含噪声;且复杂样本更易被误标。基于不完美标签的学习会导致分类器性能下降。本文提出一种新框架,通过构建最优标注者分配模型,最小化每轮循环中的最大可能噪声。同时引入新的采样方法,识别最佳查询样本,从而降低标签噪声对分类器性能的影响。实验表明,该方法显著优于多个基准方法。

原文摘要 · Abstract (English)

Active Learning (AL) has garnered significant interest across various application domains where labeling training data is costly. AL provides a framework that helps practitioners query informative samples for annotation by oracles (labelers). However, these labels often contain noise due to varying levels of labeler accuracy. Additionally, uncertain samples are more prone to receiving incorrect labels because of their complexity. Learning from imperfectly labeled data leads to an inaccurate classifier. We propose a novel AL framework to construct a robust classification model by minimizing noise levels. Our approach includes an assignment model that optimally assigns query points to labelers, aiming to minimize the maximum possible noise within each cycle. Additionally, we introduce a new sampling method to identify the best query points, reducing the impact of label noise on classifier performance. Our experiments demonstrate that our approach significantly improves classification performance compared to several benchmark methods.

主动学习标注噪声鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。