arXiv:2606.29888cs.LGcs.CV2026-06

发现视觉与语言特征方向不一致,提出分模态自编码器提升跨模态对齐

Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders

论文配图:Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
图 1 · 摘自论文原文
  • 分模态训练稀疏自编码器,保留各模态特征几何结构
  • 实验证明同概念在图像和文本中方向不同,导致模态分裂
  • 适用于需要精确控制语义的生成与检索任务

视觉-语言模型将图像与文本映射到共享嵌入空间,但这些嵌入常混杂多个语义特征,影响可解释性与可控性。尽管稀疏自编码器被用于将嵌入分解为单语义特征,其在联合嵌入空间的应用仍依赖于一个未经检验的隐含假设:语义对应的特征在跨模态间共享相同方向。本文揭示了同一概念在图像与文本模态中存在特征方向差异,称为跨模态特征异质性。我们证明该异质性是导致模态分裂的关键原因——同一概念在不同模态下激活不同潜在变量。这一发现说明仅对齐潜在激活不足以解决特征不匹配问题。为此,我们提出先分别训练模态特定的稀疏自编码器以保持各模态特征几何结构,再进行后处理对齐。该方法显著提升重建保真度,并改善跨模态检索与概念操控性能。

原文摘要 · Abstract (English)

Vision-language models map images and text into a joint embedding space. However, these embeddings often entangle multiple semantic features, which limits their interpretability and controllability. While sparse autoencoders have emerged as a useful tool for decomposing these embeddings into monosemantic features, their application to joint embedding spaces has largely relied on an implicit, untested assumption that semantically corresponding features share the same directions across modalities. In this paper, we challenge this assumption by identifying discrepancies in feature directions for the same concept across image and text modalities, a phenomenon we term cross-modal feature heterogeneity. We demonstrate that this heterogeneity is a key driver of the modality split, where a shared concept activates different latents depending on the modality. This finding further reveals why aligning latent activations alone is insufficient to resolve the underlying feature mismatch. Motivated by this observation, we propose an approach that trains modality-specific sparse autoencoders to preserve each modality's feature geometry, and then aligns corresponding features post hoc. Our method improves reconstruction fidelity and enhances performance in cross-modal retrieval and concept steering.

稀疏自编码器跨模态对齐特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。