发现视觉语言模型第一层就实现图文对齐,打破传统认知。
Representations of Text and Images Align From Layer One
- 用优化合成图像的方法,从文本概念反推视觉表示
- 层1中50%以上合成图像能识别出动物、活动或季节特征
- 无需额外数据或模型,可直观解释模型内部表征
我们发现,在基于适配器的视觉-语言模型中,多种概念的图像与文本描述在第一层就存在有意义的对齐,这与以往认为图文对齐仅出现在深层的观点相悖。我们提出一种受DeepDream启发的新型合成方法:给定如“Jupiter”这样的文本概念,提取其在某一层的概念向量,并通过优化生成与该向量对齐的图像。我们在Gemma 3模型的七层中对数百个概念进行测试,结果显示,即使在第1层,超过50%的合成图像也表现出目标概念的显著视觉特征,如动物、动作或季节。该方法提供了逐概念、逐层的直接且构造性证据,证明了图文对齐的存在。相比以往测量多模态对齐的方法,本方法简单快速,无需辅助模型或数据集,还为模型可解释性提供了新路径,可通过反向追踪图像处理组件来可视化模型的表示空间。
原文摘要 · Abstract (English)
We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the very first layer. This contradicts the established view that such image-text alignment only appears in late layers. We show this using a new synthesis-based method inspired by DeepDream: given a textual concept such as "Jupiter", we extract its concept vector at a given layer, and then use optimisation to synthesise an image whose representation aligns with that vector. We apply our approach to hundreds of concepts across seven layers in Gemma 3, and find that the synthesised images often depict salient visual features of the targeted textual concepts: for example, already at layer 1, more than 50 % of images depict recognisable features of animals, activities, or seasons. Our method thus provides direct, constructive evidence of image-text alignment on a concept-by-concept and layer-by-layer basis. Unlike previous methods for measuring multimodal alignment, our approach is simple, fast, and does not require auxiliary models or datasets. It also offers a new path towards model interpretability, by providing a way to visualise a model's representation space by backtracing through its image processing components.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。