arXiv:2509.22524cs.CV2025-09被引 2

评测视觉语言模型的配色能力,发现它们对常见颜色命名准确,但对新颜色表现差。

Color Names in Vision-Language Models

  • 用经典色彩命名实验方法测试5个主流模型的配色能力。
  • 模型在典型颜色上准确率高,但在扩展颜色集上性能大幅下降。
  • 不同模型使用不同命名策略,且英文中文占主导,适合研究人机交互者看。

颜色是人类视觉感知的基本维度,也是描述物体与场景的主要方式。随着视觉语言模型(VLMs)日益普及,理解其是否像人类一样命名颜色,对实现有效的人机交互至关重要。我们首次系统评估了多种视觉语言模型的色彩命名能力,基于经典方法使用957个颜色样本,覆盖五个代表性模型。结果显示,尽管模型在经典研究中的典型颜色上表现良好,但在扩展的非典型颜色集合上性能显著下降。我们识别出21个在所有模型中反复出现的常用颜色术语,揭示两种不同策略:受限模型主要使用基本色名,而扩展模型则采用系统性明度修饰词。跨语言分析显示,九种语言中存在严重训练不平衡,英语和中文占据主导地位,色调是颜色命名决策的主要驱动因素。消融实验表明,语言模型架构对颜色命名有显著影响,独立于视觉处理能力。

原文摘要 · Abstract (English)

Color serves as a fundamental dimension of human visual perception and a primary means of communicating about objects and scenes. As vision-language models (VLMs) become increasingly prevalent, understanding whether they name colors like humans is crucial for effective human-AI interaction. We present the first systematic evaluation of color naming capabilities across VLMs, replicating classic color naming methodologies using 957 color samples across five representative models. Our results show that while VLMs achieve high accuracy on prototypical colors from classical studies, performance drops significantly on expanded, non-prototypical color sets. We identify 21 common color terms that consistently emerge across all models, revealing two distinct approaches: constrained models using predominantly basic terms versus expansive models employing systematic lightness modifiers. Cross-linguistic analysis across nine languages demonstrates severe training imbalances favoring English and Chinese, with hue serving as the primary driver of color naming decisions. Finally, ablation studies reveal that language model architecture significantly influences color naming independent of visual processing capabilities.

视觉语言模型颜色命名人机交互多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。