arXiv:2512.02340cs.AIcs.CV2025-12被引 3

构建认知基准,诊断视觉语言模型多视角空间推理缺陷

Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective

  • 设计可系统控制视角与问题类型的认知基准,分离多视角推理与单视感知
  • 15个主流模型在跨视角对齐上表现差,信息融合阶段性能骤降超过40%
  • 通过显式/隐式分析揭示模型在推理中逐步丢失关键信息、路径不确定性上升

空间推理是人类智能的核心能力,支持三维环境中的感知、推断与规划。然而,当前视觉语言模型(VLMs)在多视角设置下难以维持几何一致性与跨视角连贯性。我们归因于缺乏细粒度的基准来隔离多视角推理与单视感知及时间因素。为此,我们提出ReMindView-Bench,一个基于认知科学的基准,用于评估VLM如何构建、对齐并维持跨互补视角的空间心智模型。该基准系统性地改变视角空间模式与查询类型,以探查空间认知的关键因素。对15个当前VLM的评估显示,其在跨视角对齐和视角转换上存在持续性失败,促使进一步分析推理过程。使用LLM-as-a-judge和自一致性提示的分阶段显式分析表明,模型在帧内感知表现良好,但在跨视角信息整合时性能急剧下降。隐式分析(包括线性探测与熵动态)进一步揭示任务相关信息逐步丢失,且正确与错误轨迹间的不确定性逐渐分离。这些结果为VLM空间推理提供了认知基础诊断,揭示了多视角空间心智模型在推理各阶段的形成、退化与失稳机制。ReMindView-Bench基准可在https://huggingface.co/datasets/Xue0823/ReMindView-Bench获取,基准构建与推理分析代码见https://github.com/pittisl/ReMindView-Bench。

原文摘要 · Abstract (English)

Spatial reasoning is a core aspect of human intelligence that allows perception, inference and planning in 3D environments. However, current vision-language models (VLMs) struggle to maintain geometric coherence and cross-view consistency for spatial reasoning in multi-view settings. We attribute this gap to the lack of fine-grained benchmarks that isolate multi-view reasoning from single-view perception and temporal factors. To address this, we present ReMindView-Bench, a cognitively grounded benchmark for evaluating how VLMs construct, align and maintain spatial mental models across complementary viewpoints. ReMindView-Bench systematically varies viewpoint spatial pattern and query type to probe key factors of spatial cognition. Evaluations of 15 current VLMs reveals consistent failures in cross-view alignment and perspective-taking in multi-view spatial reasoning, motivating deeper analysis on the reasoning process. Explicit phase-wise analysis using LLM-as-a-judge and self-consistency prompting shows that VLMs perform well on in-frame perception but degrade sharply when integrating information across views. Implicit analysis, including linear probing and entropy dynamics, further show progressive loss of task-relevant information and uncertainty separation between correct and incorrect trajectories. These results provide a cognitively grounded diagnosis of VLM spatial reasoning and reveal how multi-view spatial mental models are formed, degraded and destabilized across reasoning phases. The ReMindView-Bench benchmark is available at https://huggingface.co/datasets/Xue0823/ReMindView-Bench, and the source codes of benchmark construction and VLM reasoning analysis are available at https://github.com/pittisl/ReMindView-Bench.

空间推理多视角认知科学模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。