构建首个网页多标签识别基准,提出双分支对比学习方法提升标注噪声下的识别准确率。
Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning

- 设计双分支对比学习框架,同时捕捉实例与类别级特征及其相似性。
- 在30万张网页图像上测试,性能优于现有基线模型。
- 提供两个标准数据集,支持公平比较,推动该领域发展。
利用免费获取的网页图像训练深度学习模型可降低对昂贵人工标注的依赖。尽管网页监督学习在单标签识别中已有广泛研究,其多标签版本仍处于探索阶段,部分原因在于缺乏统一的基准和公平的评估协议。为此,我们构建了网页监督多标签识别(WS-MLR)基准,包括Web-COCO和Web-Pascal两个数据集,覆盖与MS-COCO、Pascal VOC相同的80个和20个类别,共约30万张通过类别关键词搜索获取的互联网图像。我们在统一设置下复现了代表性基线模型。进一步提出双分支多标签对比学习(DBMLCL)框架,通过联合学习类别特定的实例级与类别级表示及其相似性,实现噪声标签的识别与修正。大量实验表明,DBMLCL在该基准上显著优于现有方法。
原文摘要 · Abstract (English)
Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learning has been widely studied for single-label recognition, its multi-label counterpart remains underexplored, partly due to the lack of unified benchmarks and fair comparison protocols. To address this gap, we construct a benchmark for webly supervised multi-label recognition (WS-MLR), including Web-COCO and Web-Pascal, and re-implement representative baselines under a unified setting. The two datasets cover the same 80 and 20 categories as MS-COCO and Pascal VOC, respectively, and contain about 300 thousand images retrieved from the Internet using category-word combinations as search keywords. We further propose a Dual-Branch Multi-Label Contrastive Learning (DBMLCL) framework, which learns category-specific instance-level and category-level representations together with their similarities to identify and correct noisy labels. Extensive experiments on the benchmark demonstrate that DBMLCL achieves superior performance compared to representative baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。