用视觉错觉测试大模型的感知与推理能力,发现其表现远未达标。
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

- 以真实世界视觉错觉为诊断工具,联合评估模型感知与推理能力。
- 在自建基准上,多数大模型的推理能力被高估,表现不理想。
- 适合关注模型认知真实性与评测方法的AI研究者参考。
大型视觉语言模型具备推理能力,推动认知性能达到新高度。然而现有评估或仅关注感知,或依赖数学、编程等特定领域,缺乏对开放世界环境中感知与推理协同能力的评估。为填补这一空白,我们提出利用视觉错觉作为诊断工具来评估LVLMs。视觉错觉是人眼系统误读客观信号导致理解偏离现实的现象。我们构建了Illusion-Reasoning基准,收录真实世界的错觉图像,并包含多样化的标注问答对。基于该基准,我们发现多种主流LVLM的推理能力远低于宣称水平。本工作揭示了当前模型认知局限性,为未来优化提供方向。项目已公开于 https://github.com/zhaoliangjie55/EMNLP2026_Illusion。
原文摘要 · Abstract (English)
Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed Illusion-Reasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on Illusion-Reasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future directions for optimisation. Our project is publicly available at https://github.com/zhaoliangjie55/EMNLP2026_Illusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。