arXiv:2603.16506cs.CV2026-03被引 3

构建稀疏多视角推理基准,揭示当前模型表现仅略高于随机猜测。

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations

  • 用物理仿真生成高保真3D场景与视图元数据,支持大规模数据生成。
  • 在百万级问答对上测试,多数模型性能仅略高于随机水平。
  • 提出带视觉证据的具身思维链,提升跨数据集泛化能力。

多视角视觉推理对智能系统理解复杂环境至关重要,但现有研究多集中于单图或时序密集视频场景。现实中,系统需在无显式引导下整合有限观测,而获取大规模带精确几何与语义标注的多视角数据仍具挑战。为此,我们利用物理驱动仿真构建多样且高保真的3D场景,配备精准的每视图元数据,实现可扩展的数据生成,并保持对真实场景的迁移能力。基于此,我们提出VIEW2SPACE,一个面向稀疏多视角推理的多维基准,包含支持百万级具身问答对的可扩展分离训练划分。在此基准上,对前沿视觉语言与空间模型的全面评估显示,多视角推理仍未解决,多数模型表现仅略高于随机猜测。我们进一步探究训练是否能缩小差距:提出的具身思维链(Grounded Chain-of-Thought with Visual Evidence)在中等难度下显著提升性能,并在跨数据集评估中优于现有方法。此外,我们开展难度感知的缩放分析,涵盖模型规模、数据量、推理深度与可见性约束,结果表明,几何感知在足够可见性下可通过缩放获益,但跨稀疏视图的深层组合推理仍是根本挑战。

原文摘要 · Abstract (English)

Multi-view visual reasoning is essential for intelligent systems that must understand complex environments from sparse and discrete viewpoints, yet existing research has largely focused on single-image or temporally dense video settings. In real-world scenarios, reasoning across views requires integrating partial observations without explicit guidance, while collecting large-scale multi-view data with accurate geometric and semantic annotations remains challenging. To address this gap, we leverage physically grounded simulation to construct diverse, high-fidelity 3D scenes with precise per-view metadata, enabling scalable data generation that remains transferable to real-world settings. Based on this engine, we introduce VIEW2SPACE, a multi-dimensional benchmark for sparse multi-view reasoning, together with a scalable, disjoint training split supporting millions of grounded question-answer pairs. Using this benchmark, a comprehensive evaluation of state-of-the-art vision-language and spatial models reveals that multi-view reasoning remains largely unsolved, with most models performing only marginally above random guessing. We further investigate whether training can bridge this gap. Our proposed Grounded Chain-of-Thought with Visual Evidence substantially improves performance under moderate difficulty, and generalizes to real-world data, outperforming existing approaches in cross-dataset evaluation. We further conduct difficulty-aware scaling analyses across model size, data scale, reasoning depth, and visibility constraints, indicating that while geometric perception can benefit from scaling under sufficient visibility, deep compositional reasoning across sparse views remains a fundamental challenge.

多视角推理视觉语言模型仿真数据具身认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。