arXiv:2507.14918cs.CV2025-07

通过条件传输实现视觉语义精准对齐,提升多标签图像分类效果

Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification

  • 引入语义相关特征模块,强化标签特异性特征提取
  • 基于条件传输机制实现视觉与标签嵌入的细粒度对齐
  • 在VOC2007和MS-COCO上超越现有最优方法

多标签图像分类是机器学习中的关键任务,旨在为单张图像准确分配多个标签。现有方法常使用注意力机制或图卷积网络建模视觉表示,但受限于难以学习判别性语义感知特征,以及视觉表示与标签嵌入间缺乏细粒度对齐。为此,本文提出一种新方法SCT(Semantic-aware representation learning via Conditional Transport for Multi-Label Image Classification),引入语义相关特征学习模块,通过强调语义相关性与交互来提取判别性标签特定特征;同时设计基于条件传输的对齐机制,实现视觉-语义的精确对齐。在VOC2007和MS-COCO两个广泛使用的基准数据集上的大量实验验证了SCT的有效性,并证明其性能优于现有最先进方法。

原文摘要 · Abstract (English)

Multi-label image classification is a critical task in machine learning that aims to accurately assign multiple labels to a single image. While existing methods often utilize attention mechanisms or graph convolutional networks to model visual representations, their performance is still constrained by two critical limitations: the inability to learn discriminative semantic-aware features, and the lack of fine-grained alignment between visual representations and label embeddings. To tackle these issues in a unified framework, this paper proposes a novel approach named Semantic-aware representation learning via Conditional Transport for Multi-Label Image Classification (SCT). The proposed method introduces a semantic-related feature learning module that extracts discriminative label-specific features by emphasizing semantic relevance and interaction, along with a conditional transport-based alignment mechanism that enables precise visual-semantic alignment. Extensive experiments on two widely-used benchmark datasets, VOC2007 and MS-COCO, validate the effectiveness of SCT and demonstrate its superior performance compared to existing state-of-the-art methods.

多标签分类视觉语义对齐条件传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。