挑战CLIP图像嵌入存在内部错位的假设,发现其并非性能瓶颈。
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
- 从理论和实证两方面反驳图像-图像对齐缺失导致距离失准的说法
- 对比实验显示CLIP与纯图像训练模型在检索任务中表现相近
- 真正影响性能的是任务模糊性,而非预设的模态错位问题
近期研究认为,基于对比语言-图像训练的CLIP类模型在仅使用图像的任务中表现不佳,主因是跨模态对齐损失忽略了图像内部对齐,导致图像间距离校准不良。本文对此提出质疑:重新审视该假设的理论基础、支持指标及受影响的评估指标。理论层面,我们证明图像嵌入距离不存在所谓自由度;实证层面,发现语言-图像训练模型(CLIP、SigLIP)与图像-图像训练模型(DINO、SigLIP2)在常见内部对齐任务中的表现无显著差异。进一步实验表明,在图像检索与少样本分类任务中,解决任务模糊性才是取得最佳结果的关键,而非所谓内部错位。
原文摘要 · Abstract (English)
Recent research suggested that the embeddings produced by CLIP-like contrastive language-image training are suboptimal for image-only tasks. The main theory is that the inter-modal (language-image) alignment loss ignores intra-modal (image-image) alignment, leading to poorly calibrated distances between images. In this study, we question this intra-modal misalignment hypothesis. We reexamine its foundational theoretical argument, the indicators used to support it, and the performance metrics affected. For the theoretical argument, we demonstrate that there are no such supposed degrees of freedom for image embedding distances. For the empirical measures, our findings reveal they yield similar results for language-image trained models (CLIP, SigLIP) and image-image trained models (DINO, SigLIP2). This indicates the observed phenomena do not stem from a misalignment specific to the former. Experiments on the commonly studied intra-modal tasks retrieval and few-shot classification confirm that addressing task ambiguity, not supposed misalignment, is key for best results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。