arXiv:2510.15508cs.LGcs.NA2025-10被引 2

改进CLIP的相似度计算,让图文匹配更精准

Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity

  • 利用核空间内积捕捉信息互信息的线性结构
  • 理论证明可任意逼近最优相似度,实测性能超越标准CLIP
  • 适合关注多模态对齐理论优化的研究者

本文针对多模态对比预训练框架如CLIP中的相似度计算机制提出改进。已有理论表明,最优模态间相似度应对应于两模态间的点互信息(PMI)。然而当前CLIP及其变体未能充分利用PMI的潜在线性结构。为此,我们提出KME-CLIP,通过再生核希尔伯特空间中的内积来挖掘该线性结构。理论上证明了该方法可任意精度逼近PMI;实证表明,在多个检索与分类任务中,其整体性能优于标准CLIP。

原文摘要 · Abstract (English)

In this study, we propose an enhancement to the similarity computation mechanism in multi-modal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demonstrated that the optimal similarity metrics between paired modalities should correspond to the pointwise mutual information (PMI) between the two modalities. However, the current implementations of CLIP and its variants fail to fully utilize the underlying linear structure of PMI. We therefore propose KME-CLIP, which leverages this structure through the inner product in a reproducing kernel Hilbert space. We theoretically prove that our method can approximate PMI with arbitrary accuracy and empirically demonstrate that our approach overall outperforms the standard CLIP formulation across several retrieval and classification tasks.

CLIP多模态相似度计算理论优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。