arXiv:2412.18842cs.CVcs.LG2024-12

利用视觉语言模型提升半监督多标签学习的伪标签质量

Context-Based Semantic-Aware Alignment for Semi-Supervised Multi-Label Learning

  • 基于上下文设计标签特异性图像特征提取框架
  • 在多个基准数据集上显著提升伪标签准确率
  • 适合多标签图像分类与弱监督学习场景

由于真实世界中缺乏大量精确标注的多标签数据,半监督多标签学习(SSMLL)逐渐受到关注。预训练于大规模图文对的视觉语言模型(VLMs)蕴含丰富知识,有助于缓解标注数据稀缺问题。尽管现有基于微调VLM的方法在弱监督多标签学习中取得进展,但未能充分挖掘已标注数据信息以增强未标注数据的学习。本文提出一种基于上下文的语义感知对齐方法,通过利用VLM知识解决SSMLL问题。为处理图像中的多重语义,我们设计新型框架以提取标签特异性图像特征,实现文本特征与标签特异性图像特征间更紧凑的对齐,从而生成高质量伪标签。为增强模型对图像的整体理解,我们设计半监督上下文识别辅助任务,捕捉标签共现信息以提升特征表示。在多个基准数据集上的大量实验验证了所提方法的有效性。

原文摘要 · Abstract (English)

Due to the lack of extensive precisely-annotated multi-label data in real word, semi-supervised multi-label learning (SSMLL) has gradually gained attention. Abundant knowledge embedded in vision-language models (VLMs) pre-trained on large-scale image-text pairs could alleviate the challenge of limited labeled data under SSMLL setting.Despite existing methods based on fine-tuning VLMs have achieved advances in weakly-supervised multi-label learning, they failed to fully leverage the information from labeled data to enhance the learning of unlabeled data. In this paper, we propose a context-based semantic-aware alignment method to solve the SSMLL problem by leveraging the knowledge of VLMs. To address the challenge of handling multiple semantics within an image, we introduce a novel framework design to extract label-specific image features. This design allows us to achieve a more compact alignment between text features and label-specific image features, leading the model to generate high-quality pseudo-labels. To incorporate the model with comprehensive understanding of image, we design a semi-supervised context identification auxiliary task to enhance the feature representation by capturing co-occurrence information. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our proposed method.

多标签学习视觉语言模型半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。