利用视觉语言模型提升半监督多标签学习的伪标签质量
Context-Based Semantic-Aware Alignment for Semi-Supervised Multi-Label Learning
- 基于上下文设计标签特异性图像特征提取框架
- 在多个基准数据集上显著提升伪标签准确率
- 适合多标签图像分类与弱监督学习场景
由于真实世界中缺乏大量精确标注的多标签数据,半监督多标签学习(SSMLL)逐渐受到关注。预训练于大规模图文对的视觉语言模型(VLMs)蕴含丰富知识,有助于缓解标注数据稀缺问题。尽管现有基于微调VLM的方法在弱监督多标签学习中取得进展,但未能充分挖掘已标注数据信息以增强未标注数据的学习。本文提出一种基于上下文的语义感知对齐方法,通过利用VLM知识解决SSMLL问题。为处理图像中的多重语义,我们设计新型框架以提取标签特异性图像特征,实现文本特征与标签特异性图像特征间更紧凑的对齐,从而生成高质量伪标签。为增强模型对图像的整体理解,我们设计半监督上下文识别辅助任务,捕捉标签共现信息以提升特征表示。在多个基准数据集上的大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Due to the lack of extensive precisely-annotated multi-label data in real word, semi-supervised multi-label learning (SSMLL) has gradually gained attention. Abundant knowledge embedded in vision-language models (VLMs) pre-trained on large-scale image-text pairs could alleviate the challenge of limited labeled data under SSMLL setting.Despite existing methods based on fine-tuning VLMs have achieved advances in weakly-supervised multi-label learning, they failed to fully leverage the information from labeled data to enhance the learning of unlabeled data. In this paper, we propose a context-based semantic-aware alignment method to solve the SSMLL problem by leveraging the knowledge of VLMs. To address the challenge of handling multiple semantics within an image, we introduce a novel framework design to extract label-specific image features. This design allows us to achieve a more compact alignment between text features and label-specific image features, leading the model to generate high-quality pseudo-labels. To incorporate the model with comprehensive understanding of image, we design a semi-supervised context identification auxiliary task to enhance the feature representation by capturing co-occurrence information. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。