arXiv:2510.10055cs.CV2025-10

通过协同学习语义特征与标签恢复,提升不完整标签图像识别效果

Incomplete Multi-Label Image Recognition by Co-learning Semantic-Aware Features and Label Recovery

  • 设计统一框架,同时优化语义特征与缺失标签恢复
  • 在三个公开数据集上显著超越现有最佳方法
  • 适合处理标签不全的图像分类任务,如社交媒体图像分析

不完整标签的多标签图像识别是计算机视觉中一项重要且具有挑战性的任务,主要面临两大难题:学习语义感知特征和恢复缺失标签。本文提出一种语义感知特征与标签恢复协同学习框架(CSL),旨在统一解决这两个问题。具体而言,设计了语义相关特征学习模块,通过挖掘语义信息与标签相关性,捕捉鲁棒的语义表示;引入语义引导的特征增强模块,有效对齐视觉与语义空间,生成更具区分性的语义感知特征;最后构建协同学习框架,将语义特征学习与标签恢复融合,动态提升特征判别力并自适应推断缺失标签,形成双向强化机制。在三个广泛使用的公开数据集(MS-COCO、VOC2007、NUS-WIDE)上的大量实验表明,CSL 在不完整多标签图像识别任务上优于当前最优方法。

原文摘要 · Abstract (English)

Multi-label image recognition with incomplete labels is a challenging yet vital task in computer vision, which faces two fundamental challenges: learning semantic-aware features and recovering missing labels. In this paper, we propose a Co-learning framework for Semantic-aware features and Label recovery (CSL), designed to address both challenges in a unified learning paradigm. Specifically, we develop a semantic-related feature learning module that captures robust semantic-related representations by discovering semantic information and label correlations. Furthermore, a semantic-guided feature enhancement module is introduced to generate highly discriminative semantic-aware features by effectively aligning visual and semantic spaces. Finally, we present a collaborative learning framework that integrates semantic-aware feature learning with label recovery. This framework not only dynamically enhances the discriminability of semantic-aware features but also adaptively infers and recovers missing labels, thereby forming a mutually reinforcing mechanism between the two processes. Extensive experiments on three widely used public datasets (MS-COCO, VOC2007, and NUS-WIDE) demonstrate that CSL outperforms state-of-the-art methods for incomplete multi-label image recognition.

多标签识别语义学习标签恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。