让冻结的视觉语言模型摆脱视觉错觉,靠的是推理而非记忆。
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning

- 不训练模型,通过推理框架增强视觉理解
- 在幻觉测试中准确率显著提升,第二名表现
- 适合需要可靠视觉判断的AI系统开发者
尽管视觉语言模型在通用视觉任务中表现卓越,但在面对视觉错觉时仍表现出明显的脆弱性。这种失败常归因于捷径启发式策略,即模型更依赖语言先验和记忆原型而非直接视觉证据。本文提出结构化定性推理(SQI)框架,一种无需训练、以数据为中心的方法,旨在强化冻结的VLM的视觉定位能力。SQI通过三个模块实现:(1)公理约束注入,抑制错误度量估计与量化幻觉;(2)分层场景分解,将目标视觉特征从复杂背景干扰中分离;(3)反事实自验证,通过对抗性推理缓解确认偏误。该框架在DataCV 2026挑战赛(任务一:经典幻觉理解)中取得第二名。实验表明,SQI不仅显著提升各类幻觉情境下的准确率,且无需微调即可提供更强的可解释性诊断能力。结果证明,结构化定性对齐是构建下一代抗幻觉视觉语言系统的有效范式。
原文摘要 · Abstract (English)
While Vision-Language Models (VLMs) have achieved state-of-the-art performance in general visual tasks, their perceptual robustness remains remarkably brittle when confronted with optical illusions. These failures are often attributed to shortcut heuristics, where models prioritize linguistic priors and memorized prototypes over direct visual evidence. In this work, we propose Structured Qualitative Inference (SQI), a training-free, data-centric framework designed to fortify visual grounding in frozen VLMs. SQI addresses perceptual anomalies through three systematic modules: (1) Axiomatic Constraint Injection, which suppresses erroneous metric estimations and quantitative hallucinations; (2) Hierarchical Scene Decomposition, which decouples target visual manifolds from complex background distractors; and (3) Counterfactual Self-Verification, an adversarial reasoning step that mitigates confirmation bias. By orchestrating these qualitative constraints at inference time, SQI effectively aligns high-level linguistic reasoning with low-level visual perception. Our framework was evaluated on the DataCV 2026 Challenge (Task I: Classic Illusion Understanding), where it ranked 2nd place overall. Experimental results demonstrate that SQI not only significantly enhances accuracy across diverse illusion categories but also provides superior diagnostic interpretability without any model fine-tuning. Our success underscores the potential of structured qualitative grounding as a robust paradigm for developing next-generation, illusion-resistant vision-language systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。