arXiv:2501.11469cs.CVcs.LG2025-01AAAI被引 2

提出MASS框架,让图文匹配模型少依赖语言、多关注图像。

MASS: Overcoming Language Bias in Image-Text Matching

  • 通过构建多模态关联得分,降低模型对语言先验的依赖。
  • 在多个数据集上显著提升视觉匹配准确率,同时保持语义理解能力。
  • 无需额外训练,可直接嵌入现有视觉语言模型,适合快速部署。

预训练视觉-语言模型在多模态任务中取得了显著进展,包括图文检索。然而,图文匹配中的主要挑战之一是语言偏差:模型过度依赖语言先验,忽视视觉内容。为此,我们提出多模态关联得分(MASS)框架,减少对语言先验的依赖,从而提升图文匹配中的视觉准确性。该方法可无缝集成到现有视觉-语言模型中,无需额外训练。实验表明,MASS在不损失语言组合性理解的前提下,有效缓解了语言偏差。总体而言,MASS为提升视觉-语言模型的图文匹配性能提供了有前景的解决方案。

原文摘要 · Abstract (English)

Pretrained visual-language models have made significant advancements in multimodal tasks, including image-text retrieval. However, a major challenge in image-text matching lies in language bias, where models predominantly rely on language priors and neglect to adequately consider the visual content. We thus present Multimodal ASsociation Score (MASS), a framework that reduces the reliance on language priors for better visual accuracy in image-text matching problems. It can be seamlessly incorporated into existing visual-language models without necessitating additional training. Our experiments have shown that MASS effectively lessens language bias without losing an understanding of linguistic compositionality. Overall, MASS offers a promising solution for enhancing image-text matching performance in visual-language models.

图文匹配语言偏差多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。