测试视觉语言模型对颜色的感知与理解能力,发现现有模型仍严重不足。
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
- 构建多场景颜色基准测试,评估模型颜色感知与推理能力。
- 大模型表现更好,语言模块比视觉模块影响更大,但整体性能提升有限。
- 思维链可提升准确率与鲁棒性,颜色线索易被误导,需改进设计。
颜色在人类感知中至关重要,常为视觉推理提供关键线索。然而,当前视觉语言模型(VLMs)是否以及如何像人一样感知、理解并利用颜色尚不明确。本文提出ColorBench,一个专门用于评估VLM在颜色感知、推理与鲁棒性方面能力的综合性基准。通过精心设计的多样化真实应用场景,该基准测试模型对颜色的识别、基于颜色线索的语义推断,以及在不同颜色变换下的稳定性。我们对32个具有不同语言模型和视觉编码器的VLM进行了全面评估,发现:(i) 规模定律(模型越大越好)在ColorBench上依然成立,且语言模型的作用大于视觉编码器;(ii) 模型间性能差距较小,表明颜色理解未被充分重视;(iii) 思维链(CoT)推理虽以视觉任务为主,但仍能提升颜色理解的准确率与鲁棒性;(iv) VLM确实会利用颜色线索,但在部分任务中可能被误导。这些发现揭示了当前VLM在颜色理解上的关键局限,强调亟需提升多模态模型的人类级颜色认知能力。ColorBench可作为推动多模态人工智能向更高层次颜色理解迈进的基础工具。
原文摘要 · Abstract (English)
Color plays an important role in human perception and usually provides critical clues in visual reasoning. However, it is unclear whether and how vision-language models (VLMs) can perceive, understand, and leverage color as humans. This paper introduces ColorBench, an innovative benchmark meticulously crafted to assess the capabilities of VLMs in color understanding, including color perception, reasoning, and robustness. By curating a suite of diverse test scenarios, with grounding in real applications, ColorBench evaluates how these models perceive colors, infer meanings from color-based cues, and maintain consistent performance under varying color transformations. Through an extensive evaluation of 32 VLMs with varying language models and vision encoders, our paper reveals some undiscovered findings: (i) The scaling law (larger models are better) still holds on ColorBench, while the language model plays a more important role than the vision encoder. (ii) However, the performance gaps across models are relatively small, indicating that color understanding has been largely neglected by existing VLMs. (iii) CoT reasoning improves color understanding accuracies and robustness, though they are vision-centric tasks. (iv) Color clues are indeed leveraged by VLMs on ColorBench but they can also mislead models in some tasks. These findings highlight the critical limitations of current VLMs and underscore the need to enhance color comprehension. Our ColorBenchcan serve as a foundational tool for advancing the study of human-level color understanding of multimodal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。