arXiv:2410.16853cs.CVcs.IR2024-10被引 5

通过对齐维度信息与稀疏空间约束,提升图文匹配准确性

Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching

  • 对齐跨模态嵌入的维度信息,使相关性计算基于相似语义
  • 引入稀疏相关算法,筛选强关联空间关系,提升特征学习质量
  • 适合关注图文匹配、跨模态对齐的研究者和工程师

基于对比学习的模型在图文匹配任务中表现优异,其核心在于分析图像与文本对之间的相关性,涉及对应维度嵌入的跨模态交互。然而,不同模态的嵌入来自不同模型或模块,存在显著的模态差距。直接交互这些嵌入缺乏合理性,可能捕获错误的相关性。为此,本文提出一种新方法DIAS,从两方面弥合模态差距:(1) 对齐不同模态嵌入在对应维度上的信息表示,确保相关性计算基于相似信息的交互;(2) 引入跨模态与同模态未匹配对的空间约束,保障语义对齐的有效性。此外,提出稀疏相关算法,选择强相关空间关系,使模型学习更显著特征,避免弱相关干扰。大量实验表明,DIAS在Flickr30k和MSCOCO基准上实现4.3%–10.2%的rSum提升。

原文摘要 · Abstract (English)

Many contrastive learning based models have achieved advanced performance in image-text matching tasks. The key of these models lies in analyzing the correlation between image-text pairs, which involves cross-modal interaction of embeddings in corresponding dimensions. However, the embeddings of different modalities are from different models or modules, and there is a significant modality gap. Directly interacting such embeddings lacks rationality and may capture inaccurate correlation. Therefore, we propose a novel method called DIAS to bridge the modality gap from two aspects: (1) We align the information representation of embeddings from different modalities in corresponding dimension to ensure the correlation calculation is based on interactions of similar information. (2) The spatial constraints of inter- and intra-modalities unmatched pairs are introduced to ensure the effectiveness of semantic alignment of the model. Besides, a sparse correlation algorithm is proposed to select strong correlated spatial relationships, enabling the model to learn more significant features and avoid being misled by weak correlation. Extensive experiments demonstrate the superiority of DIAS, achieving 4.3\%-10.2\% rSum improvements on Flickr30k and MSCOCO benchmarks.

图文匹配跨模态对齐对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。