测试发现大模型难识别多视角图像中的空间矛盾
Multimodal Language Models Cannot Spot Spatial Inconsistencies

- 设计新任务:从双视角图中找出违背3D运动一致性的物体
- 顶尖多模态模型表现远低于人类,且结果波动大
- 适合关注视觉推理与物理理解的科研人员
空间一致性是视觉世界的基本属性,也是理解物理现实的关键。尽管近期进展显著,多模态大语言模型(MLLMs)在跨多视角进行3D几何推理方面仍存在困难。我们提出一项更具挑战性的任务:给定同一场景的两个视角,识别出违反3D运动一致性的物体。为此,我们提出一种简单且可扩展的方法,从多视角场景生成逼真的空间不一致图像对,实现对该能力的系统评估。实验结果表明,当前最先进的MLLMs在该任务上显著劣于人类观察者,且在不同场景属性下表现差异显著,暴露出对3D结构理解的脆弱与不完整。我们希望这些发现能推动更深入的物理世界建模方法发展。
原文摘要 · Abstract (English)
Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal large language models (MLLMs) often struggle to reason about 3D geometry across multiple views. Rather than asking models to describe scene attributes, we introduce a more challenging task: given two views of the same scene, identify the object that violates 3D motion consistency. We propose a simple and scalable method for generating realistic, spatially inconsistent image pairs from multi-view scenes, enabling systematic evaluation of this capability. Our results show that state-of-the-art MLLMs significantly underperform human observers and exhibit substantial variability across different scene attributes, revealing a fragile and incomplete understanding of 3D structure. We hope our findings underscore the need for approaches that develop a more deeply grounded understanding of the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。