arXiv:2502.04470cs.CVcs.AI2025-02被引 7

揭示CLIP模型在颜色理解上的两大缺陷:对灰白黑不敏感,且过度依赖文本

Color in Visual-Language Models: CLIP deficiencies

  • 通过合成数据实验发现CLIP对彩色图像标注准确,但对无色物体标签错误
  • 在斯特鲁普测试中验证文本主导现象,颜色标签易受文字干扰
  • 深层网络中大量神经元专用于文本,缺乏多模态颜色神经元

本文研究当前最具影响力的视觉语言模型CLIP(对比语言-图像预训练)如何编码颜色。通过在为该任务设计的合成数据集上进行实验,发现CLIP能正确为彩色视觉刺激分配颜色标签,但存在两大缺陷:(a) 对无色刺激(如白色、灰色、黑色)存在明显偏差,很少将其视为颜色;(b) 倾向于优先考虑文本信息而非其他视觉内容。通过全面的斯特鲁普效应测试证实,文本干扰对颜色标注具有显著影响。为进一步探究原因,分析了网络内部表示,发现深层网络中存在大量对文本敏感的神经元,而具备多模态颜色感知能力的神经元数量较少,这可能是理解颜色概念的关键。研究强调需改进神经网络中的颜色表征机制,以提升如CLIP等多模态模型在真实场景下的表现力与通用性。

原文摘要 · Abstract (English)

This work explores how color is encoded in CLIP (Contrastive Language-Image Pre-training) which is currently the most influential VML (Visual Language model) in Artificial Intelligence. After performing different experiments on synthetic datasets created for this task, we conclude that CLIP is able to attribute correct color labels to colored visual stimulus, but, we come across two main deficiencies: (a) a clear bias on achromatic stimuli that are poorly related to the color concept, thus white, gray and black are rarely assigned as color labels; and (b) the tendency to prioritize text over other visual information. Here we prove it is highly significant in color labelling through an exhaustive Stroop-effect test. With the aim to find the causes of these color deficiencies, we analyse the internal representation at the neuron level. We conclude that CLIP presents an important amount of neurons selective to text, specially in deepest layers of the network, and a smaller amount of multi-modal color neurons which could be the key of understanding the concept of color properly. Our investigation underscores the necessity of refining color representation mechanisms in neural networks to foster a more comprehensive comprehension of colors as humans understand them, thereby advancing the efficacy and versatility of multimodal models like CLIP in real-world scenarios.

CLIP颜色理解多模态神经机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。