arXiv:2508.20381cs.CV2025-08ICCV被引 2

解决单正样本多标签学习中的伪标签噪声问题,提升模型性能。

More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning

  • 设计新型鲁棒损失函数,有效利用多样伪标签
  • 在四个数据集上达到当前最优效果
  • 适合资源受限下高质量多标签学习场景

多标签学习是计算机视觉中一项挑战性任务,需为每张图像分配多个类别。但大规模数据集的完整标注成本高昂,难以实施,因此研究部分标注数据的学习方法成为关键。在极端情况下的单正样本多标签学习(SPML)中,每张图像仅提供一个正标签,其余标签均未标注。传统SPML方法将缺失标签视为未知或负类,易产生误判和假负例;而融合多种伪标签策略又会引入额外噪声。为此,本文提出广义伪标签鲁棒损失(GPR Loss),一种能有效利用多样化伪标签并抑制噪声的新损失函数。同时,提出简单高效的动态增强多焦点伪标签技术(DAMP)。二者共同构成自适应高效视觉-语言伪标签框架(AEVLP)。在四个基准数据集上的大量实验表明,该框架显著提升多标签分类性能,达到当前最佳水平。

原文摘要 · Abstract (English)

Multi-label learning is a challenging computer vision task that requires assigning multiple categories to each image. However, fully annotating large-scale datasets is often impractical due to high costs and effort, motivating the study of learning from partially annotated data. In the extreme case of Single Positive Multi-Label Learning (SPML), each image is provided with only one positive label, while all other labels remain unannotated. Traditional SPML methods that treat missing labels as unknown or negative tend to yield inaccuracies and false negatives, and integrating various pseudo-labeling strategies can introduce additional noise. To address these challenges, we propose the Generalized Pseudo-Label Robust Loss (GPR Loss), a novel loss function that effectively learns from diverse pseudo-labels while mitigating noise. Complementing this, we introduce a simple yet effective Dynamic Augmented Multi-focus Pseudo-labeling (DAMP) technique. Together, these contributions form the Adaptive and Efficient Vision-Language Pseudo-Labeling (AEVLP) framework. Extensive experiments on four benchmark datasets demonstrate that our framework significantly advances multi-label classification, achieving state-of-the-art results.

多标签学习伪标签视觉-语言鲁棒训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。