arXiv:2505.05721cs.CV2025-05IJCAI被引 2

通过语义空间引导的渐进式对齐,提升视觉分类中的跨模态特征匹配效果。

Semantic-Space-Intervened Diffusive Alignment for Visual Classification

  • 构建语义空间作为视觉到文本的桥梁,分两阶段实现渐进对齐。
  • 在多个数据集上优于现有方法,显著增强跨模态特征一致性。
  • 适合需要高精度跨模态对齐的视觉分类任务研究者。

跨模态对齐是提升视觉分类性能的有效方法。现有方法通常采用单步映射,利用深度神经网络将视觉特征投影以模仿文本特征分布,但因两类特征在类别样本分布和特征值范围上差异大,难以找到有效映射。为此,本文提出一种新型语义空间干预的扩散对齐方法(SeDA),将语义空间视为视觉到文本映射的桥梁,利用两类特征在分类级别上的共性信息。更重要的是,设计了双阶段扩散框架,实现两模态的渐进对齐。首先,通过扩散控制的语义学习器,基于扩散模型交互特征与视觉特征类别中心,建模视觉特征的语义空间;随后,扩散控制的语义翻译器从该语义空间学习文本特征分布。同时,渐进特征交互网络在每一步对齐中引入分步特征交互,逐步融合文本信息至映射特征中。实验结果表明,SeDA实现了更强的跨模态特征对齐,相较于现有方法,在多种场景下均取得更优性能。

原文摘要 · Abstract (English)

Cross-modal alignment is an effective approach to improving visual classification. Existing studies typically enforce a one-step mapping that uses deep neural networks to project the visual features to mimic the distribution of textual features. However, they typically face difficulties in finding such a projection due to the two modalities in both the distribution of class-wise samples and the range of their feature values. To address this issue, this paper proposes a novel Semantic-Space-Intervened Diffusive Alignment method, termed SeDA, models a semantic space as a bridge in the visual-to-textual projection, considering both types of features share the same class-level information in classification. More importantly, a bi-stage diffusion framework is developed to enable the progressive alignment between the two modalities. Specifically, SeDA first employs a Diffusion-Controlled Semantic Learner to model the semantic features space of visual features by constraining the interactive features of the diffusion model and the category centers of visual features. In the later stage of SeDA, the Diffusion-Controlled Semantic Translator focuses on learning the distribution of textual features from the semantic space. Meanwhile, the Progressive Feature Interaction Network introduces stepwise feature interactions at each alignment step, progressively integrating textual information into mapped features. Experimental results show that SeDA achieves stronger cross-modal feature alignment, leading to superior performance over existing methods across multiple scenarios.

跨模态对齐扩散模型视觉分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。