CLIP的图文嵌入呈双椭球结构,提升不确定性建模能力。
The Double-Ellipsoid Geometry of CLIP
- 发现文本与图像嵌入位于非原点对齐的线性可分椭球壳上。
- 概念越常见,越易产生假负例,导致更高不确定性。
- 用模态均值向量即可准确估算实例的共形度,适合研究对比学习几何。
对比语言-图像预训练(CLIP)在多个机器学习领域具有重要应用价值,但其嵌入空间的几何结构仍不明确。本文研究原始未归一化嵌入,发现文本和图像分别位于以线性方式可分离的椭球壳上,且这些椭球不以原点为中心。该结构有助于在对比训练中根据实例的不确定性更好地进行嵌入。数据集中频繁出现的概念会产生更多假负例,从而引发更大不确定性。本文提出一种新的共形度概念,用于衡量实例与代表性数据集中其他实例的平均余弦相似度,并证明只需计算实例与模态均值向量的余弦相似度即可准确估计。此外,我们发现CLIP的模态差距优化了图像与文本共形度分布的匹配。
原文摘要 · Abstract (English)
Contrastive Language-Image Pre-Training (CLIP) is highly instrumental in machine learning applications within a large variety of domains. We investigate the geometry of this embedding, which is still not well understood. We examine the raw unnormalized embedding and show that text and image reside on linearly separable ellipsoid shells, not centered at the origin. We explain the benefits of having this structure, allowing to better embed instances according to their uncertainty during contrastive training. Frequent concepts in the dataset yield more false negatives, inducing greater uncertainty. A new notion of conformity is introduced, which measures the average cosine similarity of an instance to any other instance within a representative data set. We show this measure can be accurately estimated by simply computing the cosine similarity to the modality mean vector. Furthermore, we find that CLIP's modality gap optimizes the matching of the conformity distributions of image and text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。