用标准色测试视觉模型能否从灰度图中还原概念信息。
Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
- 用标准色物体数据集,对比彩色与灰度图像的解码能力。
- 灰度图仍可准确解码标准色,且与物体识别相关联。
- 适合研究视觉模型概念表征的学者,尤其关注语义解码者。
视觉编码器为视觉语言模型构建图像表示。这些表示包含多少概念性而非仅直观可见的信息?我们以标准色为受控测试案例,探究视觉编码器是否能线性解码去除颜色后的图像中的标准色信息。构建了具有标准色的物体数据集,使用彩色和灰度图像探测编码器对颜色和物体身份的解码能力。结果发现,即使在灰度图像中,标准色仍可被有效解码,且与预测的物体身份紧密关联,表明存在概念性联系。将分析扩展至完整视觉语言模型(VLMs)后发现,VLM的后期训练对视觉编码器的颜色解码能力有显著影响。总体而言,标准色为追踪视觉编码器与VLM中物体级概念语义信息提供了一个可控性强的有效工具。
原文摘要 · Abstract (English)
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, information does this representation contain? We use canonical color as a controlled test case to ask whether vision encoders make canonical-color information linearly accessible, even when color is removed from the input image. We construct a dataset of objects with canonical colors, and probe vision encoders for both color and object identity using color and grayscale images. We find that canonical color remains decodable from grayscale images, and is tied to predicted object identity, indicating a conceptual link. Extending this analysis to full VLMs, we find that VLM post-training can have a surprisingly large effect on color decodability in the vision encoder. Overall, canonical color provides a usefully controllable lens for tracing object-level conceptual semantic information in vision encoders and VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。