构建开放世界空间推理基准,揭示多模态大模型依赖语言先验的缺陷
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
- 用同步立体视觉+激光雷达数据生成可量化空间问题
- 室内表现优异的模型在开放场景中性能大幅下降
- 适合研究具身智能与物理感知的学者参考
尽管多模态大语言模型在语义任务上表现优异,其空间智能——对鲁棒、具身人工智能系统至关重要的能力——仍严重不足。现有评测基准要么聚焦过于简化的定性推理,要么依赖特定室内数据,受限于缺乏带可验证度量真值的室外数据集。为此,我们构建了一个大规模基准,基于行人视角视频,配有同步立体相机、激光雷达和IMU/GPS传感器。该数据集提供精确的3D度量信息,可自动生成涵盖从定性关系到定量度量与运动学理解的多层次空间推理问题。评估显示,结构化室内基准中的性能提升在开放世界设置下消失。通过合成异常场景和盲测分析进一步证实,当前MLLMs严重依赖语言先验而非具身视觉推理。本基准为诊断这些局限性提供了原则性平台,推动物理具身空间智能的发展。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remains underdeveloped. Existing benchmarks fall short of diagnosing this limitation: they either focus on overly simplified qualitative reasoning or rely on domain-specific indoor data, constrained by the lack of outdoor datasets with verifiable metric ground truth. To bridge this gap, we introduce a large-scale benchmark built from pedestrian-perspective videos captured with synchronized stereo cameras, LiDAR, and IMU/GPS sensors. This dataset provides metrically precise 3D information, enabling the automatic generation of spatial reasoning questions that span a hierarchical spectrum--from qualitative relational reasoning to quantitative metric and kinematic understanding. Evaluations reveal that the performance gains observed in structured indoor benchmarks vanish in open-world settings. Further analysis using synthetic abnormal scenes and blinding tests confirms that current MLLMs depend heavily on linguistic priors instead of grounded visual reasoning. Our benchmark thus provides a principled platform for diagnosing these limitations and advancing physically grounded spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。