arXiv:2504.04506cs.LGcs.CV2025-04被引 5

针对标注噪声,提出新主动学习框架提升低预算下模型性能

Active Learning with a Noisy Annotator

  • 基于覆盖度的贪心策略,识别因噪声样本未覆盖区域
  • 在多个数据集上显著提升不同噪声类型下的主动学习效果
  • 适合标注质量差、预算低的视觉任务场景

主动学习旨在通过有策略地选择最具有信息量的样本以降低标注成本。然而,在仅有少量标注样本的低预算情况下,大多数主动学习方法表现不佳,尤其当标注者提供噪声标签时更为严重。现有低中预算阶段的主流方法通常聚焦于最大化已标注集对全数据集的覆盖范围。本文提出一种名为噪声感知主动采样(NAS)的新框架,将现有基于覆盖度的贪心主动学习策略扩展至处理噪声标注场景。NAS能够识别因错误代表样本导致的未覆盖区域,并支持对这些区域重新采样。同时,我们设计了一种适用于低预算的简单有效噪声过滤方法,该方法利用NAS内部机制,可在模型训练前用于清理噪声标签。在CIFAR100和ImageNet子集等多个计算机视觉基准上,NAS显著提升了标准主动学习方法在不同噪声类型与噪声率下的性能。

原文摘要 · Abstract (English)

Active Learning (AL) aims to reduce annotation costs by strategically selecting the most informative samples for labeling. However, most active learning methods struggle in the low-budget regime where only a few labeled examples are available. This issue becomes even more pronounced when annotators provide noisy labels. A common AL approach for the low- and mid-budget regimes focuses on maximizing the coverage of the labeled set across the entire dataset. We propose a novel framework called Noise-Aware Active Sampling (NAS) that extends existing greedy, coverage-based active learning strategies to handle noisy annotations. NAS identifies regions that remain uncovered due to the selection of noisy representatives and enables resampling from these areas. We introduce a simple yet effective noise filtering approach suitable for the low-budget regime, which leverages the inner mechanism of NAS and can be applied for noise filtering before model training. On multiple computer vision benchmarks, including CIFAR100 and ImageNet subsets, NAS significantly improves performance for standard active learning methods across different noise types and rates.

主动学习噪声标签视觉任务低预算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。