揭示视觉语言模型嵌入空间的稀疏线性结构与跨模态语义关联。
Interpreting the linear structure of vision-language model embedding spaces
- 用稀疏自编码器解析模型嵌入,发现概念以稀疏方向存在。
- 多数概念跨模态共激活,且方向几何对齐,支持跨模态整合。
- 释放交互式演示,可探索各模型概念空间组织规律。
视觉语言模型将图像与文本编码至联合空间,通过最小化对应图文对的距离实现对齐。为探究语言与图像在该空间中的组织方式及意义与模态的编码机制,我们对四种模型(CLIP、SigLIP、SigLIP2 和 AIMv2)的嵌入空间训练并发布稀疏自编码器(SAEs)。SAEs 将模型嵌入近似为学习到的方向(即“概念”)的稀疏线性组合。相比其他线性特征学习方法,SAEs 在重建真实嵌入的同时保持更高稀疏性。不同种子或数据饮食下重训发现:罕见特定概念易变,但高频激活的概念具有显著稳定性。尽管多数概念主要激活于单一模态,但它们几乎正交于定义模态的子空间,且无法有效区分模态,表明其编码的是跨模态语义。为此引入桥接分数(Bridge Score),量化在对齐图文输入中同时激活且空间对齐的概念对,揭示即使单模态概念也能协同支持跨模态整合。我们发布所有模型的交互式演示,供研究者探索概念空间结构。总体而言,研究揭示了视觉语言模型嵌入空间中由模态塑造却由潜在桥梁连接的稀疏线性结构,为多模态意义构建提供了新视角。
原文摘要 · Abstract (English)
Vision-language models encode images and text in a joint space, minimizing the distance between corresponding image and text pairs. How are language and images organized in this joint space, and how do the models encode meaning and modality? To investigate this, we train and release sparse autoencoders (SAEs) on the embedding spaces of four vision-language models (CLIP, SigLIP, SigLIP2, and AIMv2). SAEs approximate model embeddings as sparse linear combinations of learned directions, or "concepts". We find that, compared to other methods of linear feature learning, SAEs are better at reconstructing the real embeddings, while also able to retain the most sparsity. Retraining SAEs with different seeds or different data diet leads to two findings: the rare, specific concepts captured by the SAEs are liable to change drastically, but we also show that commonly-activating concepts are remarkably stable across runs. Interestingly, while most concepts activate primarily for one modality, we find they are not merely encoding modality per se. Many are almost orthogonal to the subspace that defines modality, and the concept directions do not function as good modality classifiers, suggesting that they encode cross-modal semantics. To quantify this bridging behavior, we introduce the Bridge Score, a metric that identifies concept pairs which are both co-activated across aligned image-text inputs and geometrically aligned in the shared space. This reveals that even single-modality concepts can collaborate to support cross-modal integration. We release interactive demos of the SAEs for all models, allowing researchers to explore the organization of the concept spaces. Overall, our findings uncover a sparse linear structure within VLM embedding spaces that is shaped by modality, yet stitched together through latent bridges, offering new insight into how multimodal meaning is constructed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。