测试视觉语言模型的推理能力,发现其主要受限于感知而非逻辑。
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
- 设计VRIQ基准,分抽象谜题与自然图像两类任务评估
- 抽象任务准确率仅28%,自然任务45%,工具辅助提升有限
- 超半数失败源于感知缺陷,形状/位置等特定感知项问题更严重
近期视觉语言模型(VLMs)的发展引发了对其非语言推理能力可靠性的疑问。为此,我们提出VRIQ(视觉推理智商)基准,用于评估和分析VLM的视觉推理能力。我们在两组任务上测试模型:抽象谜题类任务和自然图像推理任务。结果显示,在抽象谜题上性能接近随机,平均准确率约为28%;在自然任务中表现更好但依然薄弱,准确率为45%。工具增强的推理仅带来适度改进。为揭示弱点根源,我们引入诊断探针,分别针对感知与推理能力。分析表明,约56%的失败仅由感知问题导致,43%源于感知与推理共同作用,仅有1%完全由推理缺陷引起。进一步设计细粒度诊断探针,针对形状、数量、位置、三维/深度等感知类别,发现部分类别引发更多失败。本基准与分析表明,当前VLM即使使用视觉推理工具,仍无法可靠完成抽象推理,主要受限于感知能力,并为多模态系统提升视觉推理提供了系统性依据。
原文摘要 · Abstract (English)
Recent progress in Vision Language Models (VLMs) has raised the question of whether they can reliably perform nonverbal reasoning. To this end, we introduce VRIQ (Visual Reasoning IQ), a novel benchmark designed to assess and analyze the visual reasoning ability of VLMs. We evaluate models on two sets of tasks: abstract puzzle-style and natural-image reasoning tasks. We find that on abstract puzzles, performance remains near random with an average accuracy of around 28%, while natural tasks yield better but still weak results with 45% accuracy. We also find that tool-augmented reasoning demonstrates only modest improvements. To uncover the source of this weakness, we introduce diagnostic probes targeting perception and reasoning. Our analysis demonstrates that around 56% of failures arise from perception alone, 43% from both perception and reasoning, and only a mere 1% from reasoning alone. This motivates us to design fine-grained diagnostic probe questions targeting specific perception categories (e.g., shape, count, position, 3D/depth), revealing that certain categories cause more failures than others. Our benchmark and analysis establish that current VLMs, even with visual reasoning tools, remain unreliable abstract reasoners, mostly due to perception limitations, and offer a principled basis for improving visual reasoning in multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。