arXiv:2509.19070cs.CVcs.CL2025-09中稿 · the Open Science f…被引 2

用色盲测试挑战视觉语言模型,发现其易产生幻觉。

ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?

  • 设计500张类伊希哈拉图,测试模型在复杂颜色模式中识数能力
  • 9个模型在对抗性场景下识别准确率显著下降,普遍存在幻觉
  • 适合关注模型鲁棒性与真实场景可靠性的研究者使用

本文提出ColorBlindnessEval,一个新颖的基准测试,用于评估视觉语言模型(VLMs)在受伊希哈拉色盲测试启发的视觉对抗场景下的鲁棒性。数据集包含500张类伊希哈拉图像,数字范围为0至99,采用不同色彩组合,挑战VLM准确识别嵌入复杂视觉模式中的数值信息。我们使用是/否和开放性提示评估了9个VLM,并与人类参与者表现进行对比。实验揭示模型在对抗性情境下解释数字的能力存在局限,暴露出普遍的幻觉问题。这些发现强调了提升VLM在复杂视觉环境中的鲁棒性的必要性。ColorBlindnessEval可作为评估和改进VLM在高精度要求的真实应用中可靠性的有力工具。

原文摘要 · Abstract (English)

This paper presents ColorBlindnessEval, a novel benchmark designed to evaluate the robustness of Vision-Language Models (VLMs) in visually adversarial scenarios inspired by the Ishihara color blindness test. Our dataset comprises 500 Ishihara-like images featuring numbers from 0 to 99 with varying color combinations, challenging VLMs to accurately recognize numerical information embedded in complex visual patterns. We assess 9 VLMs using Yes/No and open-ended prompts and compare their performance with human participants. Our experiments reveal limitations in the models' ability to interpret numbers in adversarial contexts, highlighting prevalent hallucination issues. These findings underscore the need to improve the robustness of VLMs in complex visual environments. ColorBlindnessEval serves as a valuable tool for benchmarking and improving the reliability of VLMs in real-world applications where accuracy is critical.

视觉语言模型鲁棒性评估色盲测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。