提出新框架,让多模态模型在缺失模态时仍保持稳定表现
PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities
- 用提示词注意力机制增强分层对比学习
- 在缺模态场景下保持跨模态表示一致性
- 适合真实世界中不完整数据的多模态应用
整合自然语言与视觉信息的多模态模型显著提升了表征泛化能力,但在实际场景中,当某些模态缺失时性能急剧下降。这主要源于完整数据与缺失模态情形下表征学习不一致。现有方法通常采用较简单的生成手段处理缺失模态,但难以维持跨模态一致性,导致表现不佳。为此,我们提出PROMISE:一种面向缺失模态下鲁棒跨模态表征的提示词注意力分层对比学习框架。该框架创新性地将多模态提示学习融入分层对比学习,并设计了专用的提示注意力机制,能动态生成在特定模态缺失时仍具鲁棒性和一致性的表征,有效弥合完整与不完整数据间的表征鸿沟。在基准数据集上的大量实验及全面消融研究均表明,PROMISE优于当前最先进方法。
原文摘要 · Abstract (English)
Multimodal models integrating natural language and visual information have substantially improved generalization of representation models. However, their effectiveness significantly declines in real-world situations where certain modalities are missing or unavailable. This degradation primarily stems from inconsistent representation learning between complete multimodal data and incomplete modality scenarios. Existing approaches typically address missing modalities through relatively simplistic generation methods, yet these approaches fail to adequately preserve cross-modal consistency, leading to suboptimal performance. To overcome this limitation, we propose a novel multimodal framework named PROMISE, a PROMpting-Attentive HIerarchical ContraStive LEarning approach designed explicitly for robust cross-modal representation under conditions of missing modalities. Specifically, PROMISE innovatively incorporates multimodal prompt learning into a hierarchical contrastive learning framework, equipped with a specially designed prompt-attention mechanism. This mechanism dynamically generates robust and consistent representations for scenarios where particular modalities are absent, thereby effectively bridging the representational gap between complete and incomplete data. Extensive experiments conducted on benchmark datasets, along with comprehensive ablation studies, clearly demonstrate the superior performance of PROMISE compared to current state-of-the-art multimodal methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。