arXiv:2501.00848cs.CV2025-01被引 13

构建首个覆盖真实场景的视觉幻觉评测基准,揭示大模型在幻觉识别上的短板。

IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models

  • 设计包含1051张图像的综合幻觉数据集,融合经典与现实场景幻觉
  • 十款顶级模型在真/假判断任务中最高准确率仅80.59%,远低于人类表现
  • 发现主流模型对伪装成经典幻觉的陷阱图像易产生幻觉,开源模型反而更优

当前视觉语言模型虽在图像理解上表现优异,但在真实世界中的视觉幻觉识别上仍存在困难。现有基准多聚焦于经典认知幻觉,而这些已为先进模型所掌握,暴露出幻觉和感知能力有限等问题。为此,我们提出IllusionBench,一个涵盖经典认知幻觉与真实场景幻觉的综合性视觉幻觉数据集,包含1,051张图像、5,548个问答对及1,051条黄金文本描述,涵盖幻觉的存在性、成因与内容。我们评估了十款SOTA VLMs在真/假、多选与开放问答任务上的表现。此外,我们设计了模仿经典模式但实际不同的陷阱幻觉,凸显先进模型的幻觉问题。表现最佳的GPT-4o在真/假任务中准确率为80.59%,多选题为76.75%,仍显著落后于人类水平。在语义描述任务中,其对经典幻觉的幻觉导致陷阱幻觉得分极低,甚至低于部分开源模型。IllusionBench是目前规模最大、最全面的视觉语言模型幻觉理解评测基准。

原文摘要 · Abstract (English)

Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have been learned by state-of-the-art (SOTA) VLMs, revealing issues such as hallucinations and limited perceptual abilities. To address this gap, we introduce IllusionBench, a comprehensive visual illusion dataset that encompasses not only classic cognitive illusions but also real-world scene illusions. This dataset features 1,051 images, 5,548 question-answer pairs, and 1,051 golden text descriptions that address the presence, causes, and content of the illusions. We evaluate ten SOTA VLMs on this dataset using true-or-false, multiple-choice, and open-ended tasks. In addition to real-world illusions, we design trap illusions that resemble classical patterns but differ in reality, highlighting hallucination issues in SOTA models. The top-performing model, GPT-4o, achieves 80.59% accuracy on true-or-false tasks and 76.75% on multiple-choice questions, but still lags behind human performance. In the semantic description task, GPT-4o's hallucinations on classical illusions result in low scores for trap illusions, even falling behind some open-source models. IllusionBench is, to the best of our knowledge, the largest and most comprehensive benchmark for visual illusions in VLMs to date.

视觉幻觉大模型评测认知测试GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。