用色觉缺陷测试检验大模型能否理解不同人的视觉差异。
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
- 用伊希哈拉色盲测试作为可控刺激,评估模型对色觉差异的感知能力。
- 模型虽能描述色盲现象,但无法模拟色盲者的实际视觉体验。
- 适合关注无障碍设计与多模态系统包容性的研究者阅读。
大规模视觉语言模型(LVLM)正被应用于需要视觉推理的实际场景,如导航、教育和无障碍领域。随着能力提升,这些应用日益可行,但需考虑个体感知差异,而非假设统一的视觉体验。颜色感知是关键例子:它在视觉理解中至关重要,却因色觉缺陷而存在个体差异,而这一方面在多模态人工智能中几乎被忽视。本文通过伊希哈拉测试(Ishihara Test)检验LVLM是否能处理颜色感知差异。我们从生成结果、置信度及内部表征三方面评估模型行为,使用伊希哈拉色板作为受控刺激以揭示感知差异。尽管模型具备关于色觉缺陷的事实知识并能描述测试流程,却无法再现色盲者的真实视觉体验,而是始终采用正常色觉的默认输出。这表明当前系统缺乏表示非典型感知经验的能力,对无障碍部署和包容性多模态系统的应用提出警示。
原文摘要 · Abstract (English)
Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference. As capabilities improve, applications in navigation, education, and accessibility are becoming practical. These settings require accommodation of perceptual variation rather than assuming a uniform visual experience. Color perception illustrates this requirement: it is central to visual understanding yet varies across individuals due to Color Vision Deficiencies, an aspect largely ignored in multimodal AI. In this work, we examine whether LVLMs can account for variation in color perception using the Ishihara Test. We evaluate model behavior through generation, confidence, and internal representation, using Ishihara plates as controlled stimuli that expose perceptual differences. Although models possess factual knowledge about color vision deficiencies and can describe the test, they fail to reproduce the perceptual outcomes experienced by affected individuals and instead default to normative color perception. These results indicate that current systems lack mechanisms for representing alternative perceptual experiences, raising concerns for accessibility and inclusive deployment in multimodal settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。