arXiv:2606.08718cs.LGcs.AI2026-06中稿 · and published in t…

让模型主动重标噪声数据,提升标注效率与准确性

Deep Active Re-Labeling: Toward Noise-Resilient Annotation Efficiency

论文配图:Deep Active Re-Labeling: Toward Noise-Resilient Annotation Efficiency
图 1 · 摘自论文原文
  • 用主动学习策略识别潜在噪声样本并重标
  • 相同标注预算下,准确率显著优于传统主动学习
  • 适合标注质量不稳定、需提升数据可靠性的场景

深度主动学习(DAL)虽能降低人工标注成本,但其性能受人工标注错误制约。当高信息量样本中混入一定比例的标注噪声时,主动学习性能急剧下降,甚至劣于被动学习。本文首先分析了标注噪声对DAL的影响,提出一种应对策略:借鉴人类学习模式,将部分标注预算用于重新标注已标注数据。理论表明,只要模型具备识别潜在噪声的能力,仅重标少量样本即可有效净化训练集。为此,我们设计两种主动噪声采样策略,分别适用于不同场景,并分配预算重标这些样本。该方法赋予主动学习自我检查与修正能力。实验显示,在相同标注预算下,本方法更高效,最终获得更清洁的标注数据集。

原文摘要 · Abstract (English)

While Deep Active Learning (DAL) effectively reduces human annotation costs, its efficacy is constrained by human annotation errors. This is because the data sampled for active learning is assumed to be highly informative for training. When human annotators introduce errors into this informative data at a certain rate, the active learning performance drops significantly and, in some cases, even exhibits worse outcomes than passive learning. In this paper, we first analyze the impact of human annotation errors in the DAL setting. Then we propose a framework to address the human annotation noise problem for DAL. Informed by human learning patterns, the core idea of our proposed solution involves allocating a portion of the human annotation budget to re-annotate data that has already been labeled. Previous theoretical work suggests that when the model possesses a certain level of ability to identify potentially noisy data, even re-labeling a small fraction of the data can effectively remove noise from the active training set. To achieve this, we implement two active noise sampling strategies to detect noise under different circumstances and allocate a part of the annotation budget to re-annotate these instances. Our approach imbues active learning with a revisiting and introspective behavior. Our experiments demonstrate that, under the same annotation budget, our method is more data-efficient and yields a relatively noise-free annotation dataset in the end.

主动学习噪声鲁棒标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。