arXiv:2601.22150cs.CV2026-01被引 7

用视觉错觉测试大模型是看还是记,发现不同模型依赖不同机制。

Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions

  • 设计可控错觉框架,分离感知与记忆影响
  • 多数模型在图像反转后仍答错,显示记忆主导而非视觉感知
  • 不同模型表现差异大,提示需针对性评估

大型视觉语言模型(VLMs)在原始图像上常“正确”回答经典视觉错觉问题,但当错觉因素反转时,其答案依然不变,尽管人类能明显察觉视觉变化。这引发根本疑问:模型是基于视觉感知还是记忆模式作答?尽管已有研究观察到该现象,但成因尚不明确。本文提出VI-Probe,一种具有分级扰动和匹配对照(无错觉诱导因子)的可控视觉错觉框架,可解耦视觉感知与语言驱动的记忆回忆。不同于以往关注平均准确率的研究,我们采用极性翻转一致性、模板固定指数及归一化错觉乘数来衡量响应稳定性与敏感度。跨多个模型家族的实验表明,响应持续性源于异质性成因,并非单一机制:GPT-5表现出记忆覆盖,Claude-Opus-4.1体现感知-记忆竞争,Qwen系列则暗示视觉处理极限。结果挑战了单一成因观点,呼吁基于探测的评估,同时考察知识与对受控视觉变化的敏感度。数据与代码已公开于 https://sites.google.com/view/vi-probe/

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, even though the visual change is obvious to humans. This raises a fundamental question: do VLMs perceive visual changes or merely recall memorized patterns? While several studies have noted this phenomenon, the underlying causes remain unclear. To move from observations to systematic understanding, this paper introduces VI-Probe, a controllable visual-illusion framework with graded perturbations and matched visual controls (without illusion inducer) that disentangles visually grounded perception from language-driven recall. Unlike prior work that focuses on averaged accuracy, we measure stability and sensitivity using Polarity-Flip Consistency, Template Fixation Index, and an illusion multiplier normalized against matched controls. Experiments across different families reveal that response persistence arises from heterogeneous causes rather than a single mechanism. For instance, GPT-5 exhibits memory override, Claude-Opus-4.1 shows perception-memory competition, while Qwen variants suggest visual-processing limits. Our findings challenge single-cause views and motivate probing-based evaluation that measures both knowledge and sensitivity to controlled visual change. Data and code are available at https://sites.google.com/view/vi-probe/

视觉语言模型认知探测视觉错觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。