arXiv:2504.15037cs.LG2025-04被引 20

现有大模型难懂空间关系,需新方法突破。

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

  • 构建空间推理框架,分析现有模型缺陷
  • 指出仅靠扩大规模无法提升空间理解能力
  • 适合关注多模态模型真实世界应用的研究者

多模态大语言模型(MLLMs)在通用视觉-语言任务中表现优异,但近期研究揭示其在空间推理方面存在显著缺陷。这一短板严重限制了模型与物理世界的有效交互,制约其广泛应用。本文认为,单纯扩大模型规模和训练方法无法自然催生空间推理能力,必须对当前开发范式进行根本性调整。本文首先建立适用于MLLMs的空间推理综合框架,阐明其在现实应用中的关键作用;通过系统分析,考察从训练数据到推理机制等各组件对空间推理的影响,揭示核心局限并识别潜在改进路径。本工作旨在引导学术界关注这些被忽视的重要方向,推动实现接近人类水平的空间推理能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in general vision-language tasks. However, recent studies have exposed critical limitations in their spatial reasoning capabilities. This deficiency in spatial reasoning significantly constrains MLLMs' ability to interact effectively with the physical world, thereby limiting their broader applications. We argue that spatial reasoning capabilities will not naturally emerge from merely scaling existing architectures and training methodologies. Instead, this challenge demands dedicated attention to fundamental modifications in the current MLLM development approach. In this position paper, we first establish a comprehensive framework for spatial reasoning within the context of MLLMs. We then elaborate on its pivotal role in real-world applications. Through systematic analysis, we examine how individual components of the current methodology, from training data to reasoning mechanisms, influence spatial reasoning capabilities. This examination reveals critical limitations while simultaneously identifying promising avenues for advancement. Our work aims to direct the AI research community's attention toward these crucial yet underexplored aspects. By highlighting these challenges and opportunities, we seek to catalyze progress toward achieving human-like spatial reasoning capabilities in MLLMs.

多模态空间推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。