arXiv:2512.05137cs.CVcs.AI2025-12

测试视觉语言模型在彩色伪装图像下的表现,发现其识别能力严重受限。

ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images

  • 基于伊希哈拉色觉测试设计多类伪装图像,模拟真实干扰场景。
  • 人类与模型在细微色差下差距显著,最大错误率超60%。
  • 提出通用对比学习方法,提升模型对整体形状的恢复能力。

视觉语言模型虽在多模态理解方面取得进展,但在目标嵌入杂乱背景、需进行图底分离时仍表现不佳。为此,我们提出ChromouVQA,一个基于伊希哈拉风格彩色伪装图像的大规模多任务基准。通过扩展经典点阵图,引入多种填充几何形态,并调节色度差异、密度、尺寸、遮挡和旋转等参数,记录完整元数据以保证可复现性。该基准涵盖九项视觉问答任务,包括识别、计数、比较和空间推理。对人类与视觉语言模型的评估显示,在微弱色度对比或干扰性几何填充下存在巨大性能差距。我们还提出一种模型无关的对比学习方案,将轮廓与伪装图像对齐,有效提升全局形状恢复能力。ChromouVQA提供了一个紧凑、可控且可复现的评估框架。代码与数据集已开源:https://github.com/Chromou-VQA-Benchmark/Chromou-VQA。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have advanced multimodal understanding, yet still struggle when targets are embedded in cluttered backgrounds requiring figure-ground segregation. To address this, we introduce ChromouVQA, a large-scale, multi-task benchmark based on Ishihara-style chromatic camouflaged images. We extend classic dot plates with multiple fill geometries and vary chromatic separation, density, size, occlusion, and rotation, recording full metadata for reproducibility. The benchmark covers nine vision-question-answering tasks, including recognition, counting, comparison, and spatial reasoning. Evaluations of humans and VLMs reveal large gaps, especially under subtle chromatic contrast or disruptive geometric fills. We also propose a model-agnostic contrastive recipe aligning silhouettes with their camouflaged renderings, improving recovery of global shapes. ChromouVQA provides a compact, controlled benchmark for reproducible evaluation and extension. Code and dataset are available at https://github.com/Chromou-VQA-Benchmark/Chromou-VQA.

视觉问答多模态色觉测试模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。