arXiv:2601.18525cs.LGcs.CV2026-01被引 3

缩小模态间隙可显著提升聚类等分组任务性能。

Closing the Modality Gap Aligns Group-Wise Semantics

  • 设计新方法对齐双模态潜在空间,缓解语义结构差异。
  • 在实例级任务中提升有限,但在分组任务中效果显著提升。
  • 适用于需要语义聚合的场景,如聚类与跨模态分组。

在多模态学习中,CLIP 被视为实现跨模态共享潜在空间的主流方法,通过将相似表示拉近、相异表示推远来对齐语义。尽管基于 CLIP 的损失函数能在语义层面有效对齐模态,但所得潜在空间仍存在部分共享不足的问题,即所谓的“模态间隙”。虽然该现象的影响尚存争议,尤其因其对实例级任务(如检索)影响有限,但我们证明其在分组级任务(如聚类)中具有显著影响。为此,我们提出一种新方法,可在双模态设置下持续减小该差异,并可简单扩展至 n 模态情形。大量实验表明:尽管减少模态间隙仅带来微弱或不一致的实例级性能提升,却能显著增强分组级任务表现。这一发现可能重塑我们对模态间隙的理解,凸显其在需语义分组的任务中的关键作用。

原文摘要 · Abstract (English)

In multimodal learning, CLIP has been recognized as the \textit{de facto} method for learning a shared latent space across multiple modalities, placing similar representations close to each other and moving them away from dissimilar ones. Although CLIP-based losses effectively align modalities at the semantic level, the resulting latent spaces often remain only partially shared, revealing a structural mismatch known as the modality gap. While the necessity of addressing this phenomenon remains debated, particularly given its limited impact on instance-wise tasks (e.g., retrieval), we prove that its influence is instead strongly pronounced in group-level tasks (e.g., clustering). To support this claim, we introduce a novel method designed to consistently reduce this discrepancy in two-modal settings, with a straightforward extension to the general $n$-modal case. Through our extensive evaluation, we demonstrate our novel insight: while reducing the gap provides only marginal or inconsistent improvements in traditional instance-wise tasks, it significantly enhances group-wise tasks. These findings may reshape our understanding of the modality gap, highlighting its key role in improving performance on tasks requiring semantic grouping.

多模态语义对齐聚类潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。