arXiv:2506.05765cs.CVcs.CL2025-06被引 2

测试大模型能否分辨视觉错觉的真实与表象特征

Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?

  • 构建真假错觉数据集,区分真实与表面幻觉
  • 大模型对真假错觉回答一致,可能依赖先验知识
  • 揭示当前LVLM缺乏真正视觉理解,适合研究认知偏差

人类易受视觉错觉影响,是研究感知与认知的重要工具。受此启发,研究开始探讨大型视觉语言模型(LVLMs)是否也存在类似易感性。然而,现有研究多使用非抽象图像,且未区分实际与表观特征,导致对机器认知的评估模糊。为此,我们构建了一个视觉问答(VQA)数据集,分为真实错觉和虚假错觉两类,并配有对应控制图像。真实错觉中实际与表观特征不一致,而虚假错觉虽外观似错觉,但实际与表观特征相同。我们评估了LVLM在真实与虚假错觉任务中的表现,探究其是否能辨别实际与表观特征。结果表明,尽管模型看似能识别错觉并正确回答两类问题,但其对真实与虚假错觉的预测答案一致。这暗示其反应可能基于对错觉的先验知识,而非真正的视觉理解。数据集已公开于https://github.com/ynklab/FILM。

原文摘要 · Abstract (English)

Humans are susceptible to optical illusions, which serve as valuable tools for investigating sensory and cognitive processes. Inspired by human vision studies, research has begun exploring whether machines, such as large vision language models (LVLMs), exhibit similar susceptibilities to visual illusions. However, studies often have used non-abstract images and have not distinguished actual and apparent features, leading to ambiguous assessments of machine cognition. To address these limitations, we introduce a visual question answering (VQA) dataset, categorized into genuine and fake illusions, along with corresponding control images. Genuine illusions present discrepancies between actual and apparent features, whereas fake illusions have the same actual and apparent features even though they look illusory due to the similar geometric configuration. We evaluate the performance of LVLMs for genuine and fake illusion VQA tasks and investigate whether the models discern actual and apparent features. Our findings indicate that although LVLMs may appear to recognize illusions by correctly answering questions about both feature types, they predict the same answers for both Genuine Illusion and Fake Illusion VQA questions. This suggests that their responses might be based on prior knowledge of illusions rather than genuine visual understanding. The dataset is available at https://github.com/ynklab/FILM

视觉语言模型视觉错觉认知评估VQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。