构建首个覆盖真实场景的视觉幻觉评测基准,揭示大模型在幻觉识别上的短板。
IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models
- 设计包含1051张图像的综合幻觉数据集,融合经典与现实场景幻觉
- 十款顶级模型在真/假判断任务中最高准确率仅80.59%,远低于人类表现
- 发现主流模型对伪装成经典幻觉的陷阱图像易产生幻觉,开源模型反而更优
当前视觉语言模型虽在图像理解上表现优异,但在真实世界中的视觉幻觉识别上仍存在困难。现有基准多聚焦于经典认知幻觉,而这些已为先进模型所掌握,暴露出幻觉和感知能力有限等问题。为此,我们提出IllusionBench,一个涵盖经典认知幻觉与真实场景幻觉的综合性视觉幻觉数据集,包含1,051张图像、5,548个问答对及1,051条黄金文本描述,涵盖幻觉的存在性、成因与内容。我们评估了十款SOTA VLMs在真/假、多选与开放问答任务上的表现。此外,我们设计了模仿经典模式但实际不同的陷阱幻觉,凸显先进模型的幻觉问题。表现最佳的GPT-4o在真/假任务中准确率为80.59%,多选题为76.75%,仍显著落后于人类水平。在语义描述任务中,其对经典幻觉的幻觉导致陷阱幻觉得分极低,甚至低于部分开源模型。IllusionBench是目前规模最大、最全面的视觉语言模型幻觉理解评测基准。
原文摘要 · Abstract (English)
Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have been learned by state-of-the-art (SOTA) VLMs, revealing issues such as hallucinations and limited perceptual abilities. To address this gap, we introduce IllusionBench, a comprehensive visual illusion dataset that encompasses not only classic cognitive illusions but also real-world scene illusions. This dataset features 1,051 images, 5,548 question-answer pairs, and 1,051 golden text descriptions that address the presence, causes, and content of the illusions. We evaluate ten SOTA VLMs on this dataset using true-or-false, multiple-choice, and open-ended tasks. In addition to real-world illusions, we design trap illusions that resemble classical patterns but differ in reality, highlighting hallucination issues in SOTA models. The top-performing model, GPT-4o, achieves 80.59% accuracy on true-or-false tasks and 76.75% on multiple-choice questions, but still lags behind human performance. In the semantic description task, GPT-4o's hallucinations on classical illusions result in low scores for trap illusions, even falling behind some open-source models. IllusionBench is, to the best of our knowledge, the largest and most comprehensive benchmark for visual illusions in VLMs to date.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。