剖析多模态大模型空间理解短板,从数据到结构给出系统性答案。
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
- 构建多视角空间理解基准,跨单图、多图、视频三场景评估
- 数据量增加后性能快速饱和,想象类任务上限仍低
- 视觉编码器的位置编码比语言模型更关键,建议优化架构设计
空间理解对多模态大模型在具身环境中的感知、推理与规划至关重要。尽管近年有进展,现有研究仍显示其在空间理解上存在明显短板。但多数研究局限于单一场景(如单图或视频),缺乏系统性评估。本文从数据与架构双角度,针对单图、多图、视频三类典型场景展开系统分析。提出名为 MulSeT(Multi-view Spatial Understanding Tasks)的基准,并设计系列实验考察 MLLMs 的空间推理能力。从数据视角看,随着训练数据增加,空间理解性能迅速收敛,且上限较低,尤其在需空间想象的任务中表现不佳,表明单纯扩大数据无法提升效果。从架构视角看,空间理解更依赖视觉编码器中的位置编码,而非语言模型,且在级联与原生架构中均成立。进一步探索了推理注入策略,并展望通过架构改进优化空间理解。研究揭示了当前 MLLMs 的局限性,为未来通过数据扩展与架构调优提升空间推理能力提供新方向。
原文摘要 · Abstract (English)
Spatial understanding is essential for Multimodal Large Language Models (MLLMs) to support perception, reasoning, and planning in embodied environments. Despite recent progress, existing studies reveal that MLLMs still struggle with spatial understanding. However, existing research lacks a comprehensive and systematic evaluation of these limitations, often restricted to isolated scenarios, such as single-view or video. In this work, we present a systematic analysis of spatial understanding from both data and architectural perspectives across three representative scenarios: single-view, multi-view, and video. We propose a benchmark named MulSeT (Multi-view Spatial Understanding Tasks), and design a series of experiments to analyze the spatial reasoning capabilities of MLLMs. From the data perspective, the performance of spatial understanding converges quickly as the training data increases, and the upper bound is relatively low, especially for tasks that require spatial imagination. This indicates that merely expanding training data is insufficient to achieve satisfactory performance. From the architectural perspective, we find that spatial understanding relies more heavily on the positional encoding within the visual encoder than within the language model, in both cascaded and native MLLMs. Moreover, we explore reasoning injection and envision future improvements through architectural design to optimize spatial understanding. These insights shed light on the limitations of current MLLMs and suggest new directions for improving spatial reasoning capabilities through data scaling and architectural tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。