通过能量一致性揭示视觉语言模型的隐藏几何结构
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
- 基于跨模态冗余假设,设计稀疏自编码器保持能量一致
- 发现双模态原子承载全部对齐信号,单模态原子导致模态差异
- 移除单模态成分可消除差距且不损失性能,适合模型可解释性研究
视觉语言模型(VLMs)在图像与文本对齐上表现卓越,但其共享嵌入空间的几何结构仍不清晰。本文从等能量假设出发,利用跨模态冗余:真正共享的概念在不同模态中应具有相同的平均能量。为此提出对齐稀疏自编码器(SAE),在训练中鼓励能量一致性同时保持重建能力。实验表明该归纳偏置改变解空间却不损害重建效果,生成可用于几何分析的表示。在控制数据上的验证显示,当等能量成立时对齐性能提升,否则无变化。应用于基础VLMs后发现:(i) 稀疏双模态原子承载全部跨模态对齐信号;(ii) 单模态原子作为模态特异性偏差,完全解释模态差距;(iii) 移除单模态原子可消除差距而不影响性能;(iv) 仅在双模态子空间内进行向量运算能实现分布内编辑并提升检索效果。这些发现表明,合适的归纳偏置可在保持模型保真度的同时,使潜在几何结构可解释且可操作。
原文摘要 · Abstract (English)
Vision-language models (VLMs) align images and text with remarkable success, yet the geometry of their shared embedding space remains poorly understood. To probe this geometry, we begin from the Iso-Energy Assumption, which exploits cross-modal redundancy: a concept that is truly shared should exhibit the same average energy across modalities. We operationalize this assumption with an Aligned Sparse Autoencoder (SAE) that encourages energy consistency during training while preserving reconstruction. We find that this inductive bias changes the SAE solution without harming reconstruction, giving us a representation that serves as a tool for geometric analysis. Sanity checks on controlled data with known ground truth confirm that alignment improves when Iso-Energy holds and remains neutral when it does not. Applied to foundational VLMs, our framework reveals a clear structure with practical consequences: (i) sparse bimodal atoms carry the entire cross-modal alignment signal; (ii) unimodal atoms act as modality-specific biases and fully explain the modality gap; (iii) removing unimodal atoms collapses the gap without harming performance; (iv) restricting vector arithmetic to the bimodal subspace yields in-distribution edits and improved retrieval. These findings suggest that the right inductive bias can both preserve model fidelity and render the latent geometry interpretable and actionable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。