arXiv:2602.07311cs.CVcs.AI2026-02

统一视觉语言稀疏编码,实现跨模态可解释概念发现

LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept Discovery

  • 构建共享稀疏码字典,对齐图像块与文本词元表征
  • 无需标注即可实现跨模态神经元对应,提升解释性
  • 自动聚类解析词典,支持动作、属性等抽象概念识别

稀疏自编码器(SAEs)为不同表示空间提供可比的解释路径,但现有方法按模态独立训练,导致特征不可直接理解且解释无法跨域迁移。本文提出LUCID(Learning Unified vision-language sparse Codes for Interpretable concept Discovery),一种统一的视觉-语言稀疏自编码器,学习图像块与文本词元的共享潜在字典,同时保留模态特有容量。通过引入无需标注的最优传输匹配目标实现特征对齐。LUCID生成可解释的共享特征,支持像素级定位,建立跨模态神经元对应,并缓解基于相似度评估中的概念聚类问题。基于对齐特性,我们开发了无需人工观察的自动化字典解释流程。分析显示,共享特征捕捉了超越物体的多样语义类别,包括动作、属性与抽象概念,展现全面的可解释多模态表征能力。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) offer a natural path toward comparable explanations across different representation spaces. However, current SAEs are trained per modality, producing dictionaries whose features are not directly understandable and whose explanations do not transfer across domains. In this study, we introduce LUCID (Learning Unified vision-language sparse Codes for Interpretable concept Discovery), a unified vision-language sparse autoencoder that learns a shared latent dictionary for image patch and text token representations, while reserving private capacity for modality-specific details. We achieve feature alignment by coupling the shared codes with a learned optimal transport matching objective without the need of labeling. LUCID yields interpretable shared features that support patch-level grounding, establish cross-modal neuron correspondence, and enhance robustness against the concept clustering problem in similarity-based evaluation. Leveraging the alignment properties, we develop an automated dictionary interpretation pipeline based on term clustering without manual observations. Our analysis reveals that LUCID's shared features capture diverse semantic categories beyond objects, including actions, attributes, and abstract concepts, demonstrating a comprehensive approach to interpretable multimodal representations.

稀疏编码多模态可解释性概念发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。