arXiv:2602.01816cs.CV2026-02被引 12

测试大模型在视觉错觉和异常上的表现,发现其远不如人类可靠。

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

  • 构建六类错觉与异常数据集,用人工审核确保质量
  • 超20个顶尖模型测试显示普遍脆弱,尤其对错觉无抵抗力
  • 即使使用思维链也难逃幻觉,适合研究模型鲁棒性者参考

多模态大语言模型在常规视觉语言任务上已达到甚至超越人类水平,但这些评估多基于分布内数据,未检验模型在违背常识前提下的鲁棒性。为此,我们提出VIA-Bench,一个挑战性基准,用于探测模型在视觉错觉与异常场景中的表现。该基准包含六类核心内容:颜色错觉、运动错觉、格式塔错觉、几何与空间错觉、一般视觉错觉及视觉异常。通过人机协同严格审核,构建了超过1000个高质量问答对,需精细视觉推理。对超过20个前沿多模态模型(含闭源、开源及增强推理模型)的全面评估揭示显著缺陷。特别地,思维链(CoT)推理几乎无法提升鲁棒性,常产生“脆性幻觉”,即模型逻辑在错觉刺激下崩溃。结果表明机器与人类感知存在根本差异,解决此类感知瓶颈对推动通用人工智能至关重要。基准数据与代码将公开。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. However, these evaluations typically rely on standard in-distribution data, leaving the robustness of MLLMs largely unexamined when faced with scenarios that defy common-sense priors. To address this gap, we introduce VIA-Bench, a challenging benchmark designed to probe model performance on visual illusions and anomalies. It includes six core categories: color illusions, motion illusions, gestalt illusions, geometric and spatial illusions, general visual illusions, and visual anomalies. Through careful human-in-the-loop review, we construct over 1K high-quality question-answer pairs that require nuanced visual reasoning. Extensive evaluation of over 20 state-of-the-art MLLMs, including proprietary, open-source, and reasoning-enhanced models, uncovers significant vulnerabilities. Notably, we find that Chain-of-Thought (CoT) reasoning offers negligible robustness, often yielding ``brittle mirages'' where the model's logic collapses under illusory stimuli. Our findings reveal a fundamental divergence between machine and human perception, suggesting that resolving such perceptual bottlenecks is critical for the advancement of artificial general intelligence. The benchmark data and code will be released.

多模态视觉错觉模型评测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。