解决视觉语言检索中的模态不平衡问题,提升跨模态匹配效果。
Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation
- 通过结构感知蒸馏实现多粒度跨模态匹配,保留数据内在结构。
- 在多个数据集上显著提升跨模态检索性能,同时增强单模态检索能力。
- 适合关注跨模态对齐与结构保持的研究者和应用开发者。
视觉语言检索旨在基于一模态查询在另一模态中搜索相似实例,核心目标是学习跨模态匹配表示。然而,实际中噪声干扰与模态信息不足常导致模态不平衡,影响匹配效果。本文首次揭示:模态不平衡下,理想的跨模态匹配通常无法实现最优检索性能。此时,共同空间中实例的结构会受到破坏,挑战跨模态相似性度量。为此,我们强调有意义的结构保持匹配的重要性,提出一种简单有效的重平衡方法,通过学习结构保持的匹配表示来解决该问题。具体地,设计了一种新颖的多粒度跨模态匹配机制,结合跨模态匹配损失与结构感知蒸馏。前者约束实例级匹配,后者通过关系匹配正则化学习到的匹配表示与模态内表示之间的几何一致性。在多个数据集上的大量实验验证了该方法在跨模态检索上的优越性能,且相较基线模型同时提升了单模态检索能力。
原文摘要 · Abstract (English)
Vision-language retrieval aims to search for similar instances in one modality based on queries from another modality. The primary objective is to learn cross-modal matching representations in a latent common space. Actually, the assumption underlying cross-modal matching is modal balance, where each modality contains sufficient information to represent the others. However, noise interference and modality insufficiency often lead to modal imbalance, making it a common phenomenon in practice. The impact of imbalance on retrieval performance remains an open question. In this paper, we first demonstrate that ultimate cross-modal matching is generally sub-optimal for cross-modal retrieval when imbalanced modalities exist. The structure of instances in the common space is inherently influenced when facing imbalanced modalities, posing a challenge to cross-modal similarity measurement. To address this issue, we emphasize the importance of meaningful structure-preserved matching. Accordingly, we propose a simple yet effective method to rebalance cross-modal matching by learning structure-preserved matching representations. Specifically, we design a novel multi-granularity cross-modal matching that incorporates structure-aware distillation alongside the cross-modal matching loss. While the cross-modal matching loss constraints instance-level matching, the structure-aware distillation further regularizes the geometric consistency between learned matching representations and intra-modal representations through the developed relational matching. Extensive experiments on different datasets affirm the superior cross-modal retrieval performance of our approach, simultaneously enhancing single-modal retrieval capabilities compared to the baseline models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。