跨模态比较视觉、文本与多模态编码器的共享概念,揭示预训练对特征一致性的影响。
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
- 提出新指标量化不同模型间SAE特征的共享程度
- 21个编码器在双规模和通用/专用数据集上验证,发现视觉特有特征可被文本编码器共享
- 适用于研究多模态预训练如何塑造共性表征,适合模型对比与解释性研究者
稀疏自编码器(SAEs)已成为从神经网络激活中提取人类可理解特征的强大工具。以往研究基于SAE提取的特征比较不同模型,但这些比较局限于同一模态内。本文提出一种新型指标,实现跨模态模型在SAE特征上的定量比较,并用于分析视觉、文本及多模态编码器。我们还提出量化各类模型间个体特征共享度的方法。借助这两项新工具,我们在三种类型共21个编码器上开展研究,涵盖两种显著不同的规模,并考虑通用和领域特定数据集。结果重新审视了在多模态预训练背景下先前的研究,量化了各类模型共享表示的程度。结果显示,仅在视觉语言模型(VLMs)中出现的视觉特征会与文本编码器共享,凸显文本预训练的深远影响。代码已开源:https://github.com/CEA-LIST/SAEshareConcepts。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) have emerged as a powerful technique for extracting human-interpretable features from neural networks activations. Previous works compared different models based on SAE-derived features but those comparisons have been restricted to models within the same modality. We propose a novel indicator allowing quantitative comparison of models across SAE features, and use it to conduct a comparative study of visual, textual and multimodal encoders. We also propose to quantify the Comparative Sharedness of individual features between different classes of models. With these two new tools, we conduct several studies on 21 encoders of the three types, with two significantly different sizes, and considering generalist and domain specific datasets. The results allow to revisit previous studies at the light of encoders trained in a multimodal context and to quantify to which extent all these models share some representations or features. They also suggest that visual features that are specific to VLMs among vision encoders are shared with text encoders, highlighting the impact of text pretraining. The code is available at https://github.com/CEA-LIST/SAEshareConcepts
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。