探究视觉语言模型如何理解图像中的关系,发现其依赖语言线索而非真正理解视觉关系。
Investigating Relational Reasoning in VLMs

- 构建几何形状数据集,精准测试模型对视觉关系的理解能力
- 发现模型在推理中同时使用真实视觉分析与语言线索捷径
- 适合关注模型可解释性与认知机制的研究者
视觉语言模型(VLMs)在视觉推理任务中表现优异,但其是否真正理解视觉关系仍不明确,可能仅依赖语言提示或先验知识。为此,我们采用Qwen3-VL-4B(Bai et al., 2025)这一现代VLM,分析视觉信息在不同深度的编码方式。为此,我们设计了一个由简单几何形状构成的合成数据集,用于可控分析,并构建特定查询以精确测试语言提示的影响。此外,数据集还被修改以检验模型对视觉证据的因果依赖性。结果表明,当前的VLM同时结合了真实的视觉推理与主要基于语言线索的捷径策略。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we use the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths. For this, we propose a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues. Furthermore, the dataset is modified to test causal reliance on visual evidence. Our results show that current VLMs combine genuine visual reasoning with shortcut strategies primarily rooted in language cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。