通过认知框架解析视觉模型如何将感知与推理结合,发现解耦感知和推理可显著提升性能。
A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs
- 设计三类认知模拟评估范式,模拟人类解题策略。
- 在多个基准上达到新SOTA,尤其在多图推理任务中表现突出。
- 证明感知瓶颈限制推理,解耦感知与推理是未来方向。
人工智能中一个根本挑战是理解复杂模型(如视觉语言模型)进行视觉推理时的认知机制。这些模型如何将视觉感知与抽象思维结合,特别是在跨多图或需要细粒度组合理解的任务中?本文受认知科学启发,提出一种结构化评估框架,采用多种视觉推理任务——邦加德问题(Bongard Problems, BPs)和Winoground,以剖析视觉语言模型中的感知-推理界面。我们设计三种评估范式,对应人类解题策略:直接视觉规则学习(DVRL;整体处理)、演绎规则学习(DRL;规则提取与应用)、组件分析(CA;通过任务无关文本描述进行分解分析)。这些范式系统性地改变认知负荷并探测处理阶段。值得注意的是,CA可对单图架构实现多图推理评估,并通过文本描述将推理与感知分离。应用该框架,我们发现,借助强大语言模型对独立生成的丰富描述进行推理,CA在挑战性基准(包括Bongard-OpenWorld、Bongard-HOI和Winoground)上达到新状态(SOTA)。消融实验表明,当感知挑战被缓解时,推理性能显著提升,揭示了关键的感知瓶颈。该框架提供了一种有价值的诊断工具,并提示:通过丰富、任务无关的描述解耦感知与推理,是实现鲁棒且通用视觉智能的有前景方向。
原文摘要 · Abstract (English)
A fundamental challenge in artificial intelligence involves understanding the cognitive mechanisms underlying visual reasoning in sophisticated models like Vision-Language Models (VLMs). How do these models integrate visual perception with abstract thought, especially when reasoning across multiple images or requiring fine-grained compositional understanding? Drawing inspiration from cognitive science, this paper introduces a structured evaluation framework using diverse visual reasoning tasks-Bongard Problems (BPs) and Winoground-to dissect the perception-reasoning interface in VLMs. We propose three distinct evaluation paradigms, mirroring human problem-solving strategies: Direct Visual Rule Learning (DVRL; holistic processing), Deductive Rule Learning (DRL; rule extraction and application), and Componential Analysis (CA; analytical decomposition via task-agnostic textual descriptions). These paradigms systematically vary cognitive load and probe processing stages. Notably, CA enables multi-image reasoning evaluation even for single-image architectures and isolates reasoning from perception by operating on textual descriptions. Applying this framework, we demonstrate that CA, leveraging powerful language models for reasoning over rich, independently generated descriptions, achieves new state-of-the-art (SOTA) performance on challenging benchmarks including Bongard-OpenWorld, Bongard-HOI, and Winoground. Ablation studies confirm reasoning improves significantly when perceptual challenges are mitigated, revealing a critical perception bottleneck. Our framework provides a valuable diagnostic tool and suggests that decoupling perception (via rich, task-agnostic description) from reasoning is a promising direction for robust and general visual intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。