分析视觉语言模型的认知短板,发现其推理能力受限于空间理解与注意力机制。
Caption This, Reason That: VLMs Caught in the Middle
- 从感知、注意、记忆三方面测试先进模型表现
- 小模型经微调后核心认知能力显著提升
- 生成文本描述可改善视觉推理,提示需增强链式思维
近年来,视觉语言模型(VLMs)在视觉理解方面取得显著进展,但在计数或关系推理等具体任务上仍落后于人类。为揭示其根本局限,我们借鉴认知科学方法,沿感知、注意、记忆三大认知维度评估主流VLMs(包括GPT-4o)。结果表明:尽管先进模型在类别识别等任务接近人类上限,但在需要空间理解或选择性注意的任务中仍存在明显差距。通过视觉-文本解耦分析发现,难以进行直接视觉推理的模型,在基于自身生成的文本描述进行推理时表现显著提升。这说明即使性能超越人类的模型,仍亟需强化链式思维(CoT)能力。此外,针对复合视觉推理任务进行微调,能有效提升小型VLM的核心认知能力。虽该改进对复杂、分布外基准提升有限,但本研究数据集上的表现与其它基准强相关。工作系统揭示了多任务并行感知与推理的关键瓶颈,并提出简单有效的优化路径。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years. Yet, they still lag behind human capabilities in specific visual tasks such as counting or relational reasoning. To understand the underlying limitations, we adopt methodologies from cognitive science, analyzing VLM performance along core cognitive axes: Perception, Attention, and Memory. Using a suite of tasks targeting these abilities, we evaluate state-of-the-art VLMs, including GPT-4o. Our analysis reveals distinct cognitive profiles: while advanced models approach ceiling performance on some tasks (e.g. category identification), a significant gap persists, particularly in tasks requiring spatial understanding or selective attention. Investigating the source of these failures and potential methods for improvement, we employ a vision-text decoupling analysis, finding that models struggling with direct visual reasoning show marked improvement when reasoning over their own generated text captions. These experiments reveal a strong need for improved VLM Chain-of-Thought (CoT) abilities, even in models that consistently exceed human performance. Furthermore, we demonstrate the potential of targeted fine-tuning on composite visual reasoning tasks and show that fine-tuning smaller VLMs substantially improves core cognitive abilities. While this improvement does not translate to large enhancements on challenging, out-of-distribution benchmarks, we show broadly that VLM performance on our datasets strongly correlates with performance on these other benchmarks. Our work provides a detailed analysis of VLM cognitive strengths and weaknesses and identifies key bottlenecks in simultaneous perception and reasoning while also providing an effective and simple solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。