发现视觉大模型常装懂,看似正确实则幻觉。
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
- 用三层诊断框架拆解模型是否真看懂图像
- 72.9%样本出现视觉拍马屁,内部证据在却答错
- 模型越大会更依赖幻觉,但可不训练就提准
当视觉语言模型(VLM)回答正确时,它们是否真正依赖视觉信息?我们提出三层次诊断框架,包含潜空间异常检测、视觉必要性评分和竞争性评分三个每样本指标,以分离感知、依赖与对齐失败。在9个VLM和9000个样本上,通过反事实盲区、噪声和冲突干预测试,72.9%的样本表现出视觉拍马屁现象——即内部视觉证据仍存在,但模型却生成了幻觉答案;而零样本显示稳健拒绝,表明当前对齐训练已消除拒绝作为解码结果。在Qwen-VL系列中,规模扩大(包括同代和跨代)虽单调降低语言捷径,却加剧视觉拍马屁,说明仅靠规模和新后训练无法解决视觉接地问题。诊断得分还可实现无需训练的筛选预测策略,在50%覆盖下最高提升9.5个百分点准确率。
原文摘要 · Abstract (English)
When VLMs answer correctly, do they genuinely rely on visual information? We introduce a Tri-Layer Diagnostic Framework with three per-sample metrics: Latent Anomaly Detection, Visual Necessity Score, and Competition Score, which disentangle perception, dependency, and alignment failures. Across 9 VLMs and 9,000 model-sample pairs under counterfactual blind, noise, and conflict interventions, 72.9% of samples exhibit Visual Sycophancy, a Split Beliefs pattern in which internal evidence is preserved yet a hallucinated answer is decoded, while zero samples show Robust Refusal, indicating that current alignment training has eliminated refusal as a decoding outcome. Scaling within the Qwen-VL family, both within- and across-generation, monotonically reduces Language Shortcuts but amplifies Visual Sycophancy, showing that scale and newer post-training alone cannot resolve the grounding problem. Diagnostic scores further enable a training-free selective-prediction strategy yielding up to +9.5 percentage points accuracy at 50% coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。