用类别概率生成对齐多模态表示,提升跨域泛化能力
Generative Modeling of Class Probability for Multi-Modal Representation Learning
- 将类别概率作为锚点,通过提示生成跨模态分布
- 在4个基准数据集上超越现有方法,跨域表现更优
- 适合需要强泛化能力的多模态学习场景
多模态理解在人工智能中至关重要,使模型能联合解析来自不同模态的输入。然而,传统对比学习常因模态差异导致对齐偏差。本文提出一种新型类别锚点对齐方法,利用类别概率分布进行多模态表征学习。所提方法CALM(Class-anchor-ALigned generative Modeling)将类别锚点编码为提示,生成并对齐各模态的类别概率分布,实现更有效的对齐。此外,引入跨模态概率变分自编码器以建模对齐中的不确定性,增强捕捉模态间深层关系与数据变化的能力。在4个基准数据集上的大量实验表明,该方法显著优于当前最先进方法,尤其在跨域评估中表现突出,体现出更强的泛化能力。
原文摘要 · Abstract (English)
Multi-modal understanding plays a crucial role in artificial intelligence by enabling models to jointly interpret inputs from different modalities. However, conventional approaches such as contrastive learning often struggle with modality discrepancies, leading to potential misalignments. In this paper, we propose a novel class anchor alignment approach that leverages class probability distributions for multi-modal representation learning. Our method, Class-anchor-ALigned generative Modeling (CALM), encodes class anchors as prompts to generate and align class probability distributions for each modality, enabling more effective alignment. Furthermore, we introduce a cross-modal probabilistic variational autoencoder to model uncertainty in the alignment, enhancing the ability to capture deeper relationships between modalities and data variations. Extensive experiments on four benchmark datasets demonstrate that our approach significantly outperforms state-of-the-art methods, especially in out-of-domain evaluations. This highlights its superior generalization capabilities in multi-modal representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。