发现视觉语言模型的嵌入空间存在164维共享噪声,可安全移除而不影响性能。
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers

- 通过协方差谱分解分离语义信号与共享噪声
- 164维噪声在不同数据子集上具有强不变性
- 剪除噪声维度后下游任务性能保持或提升
对比预训练的视觉-语言模型(VLMs)虽是强大的特征提取器,但其共享潜在空间易受结构异常影响,成为非语义、多模态噪声的存储库。为解决此问题,我们采用协方差矩阵的谱分解,将VLM潜在空间分解为多模态语义信号分量和共享噪声子空间。观察发现,该噪声几何结构在不同数据子集间表现出强子群不变性。关键的是,剪除这些共享噪声维度主要无害,甚至能维持或提升下游任务性能。本工作揭示了现代VLM表示结构的新机制,表明其潜在几何的相当大比例由共享的、架构级噪声而非任务相关语义单独决定。
原文摘要 · Abstract (English)
Contrastively pre-trained Vision-Language Models (VLMs) serve as powerful feature extractors. Yet, their shared latent spaces are prone to structural anomalies and act as repositories for non-semantic, multi-modal noise. To address this phenomenon, we employ spectral decomposition of covariance matrices to decompose the VLM latent space into a multi-modal semantic signal component and a shared noise subspace. We observe that this noise geometry exhibits strong subgroup invariance across distinct data subsets. Crucially, pruning these shared noise dimensions is mainly harmless, preserving or actively improving downstream task performance. By isolating true semantic signals from artifactual noise, this work provides new mechanistic insights into the representational structure of modern VLMs, suggesting that a substantial fraction of their latent geometry is governed by shared, architecture-level noise rather than task-relevant semantics alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。