arXiv:2510.25760cs.CV2025-10综述被引 23

综述大模型时代的多模态空间推理,提供任务分类与开源评测基准。

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

  • 系统梳理视觉、声音等多模态信息下的空间推理方法
  • 构建涵盖2D/3D场景理解的开放评测基准集
  • 适合研究多模态智能、具身AI的学者与工程师参考

人类具备通过视觉、听觉等多模态感知理解空间的能力。大型多模态推理模型通过学习感知与推理,展现出在各类空间任务中的优异表现。然而,针对这些模型的系统性综述和公开评测基准仍较匮乏。本文全面回顾了大模型时代多模态空间推理任务,对近期多模态大语言模型(MLLMs)进展进行分类,并引入开源评测基准以支持评估。从经典二维任务出发,涵盖空间关系推理、场景与布局理解,以及三维空间中的视觉问答与定位任务;同时综述具身AI进展,包括视觉-语言导航与动作建模。此外,还探讨音频、第一人称视频等新兴模态如何借助新传感器推动空间理解。本综述为该领域奠定基础,并提供深入洞察。最新资料、代码及基准实现详见 https://github.com/zhengxuJosh/Awesome-Spatial-Reasoning。

原文摘要 · Abstract (English)

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive and reason, showing promising performance across diverse spatial tasks. However, systematic reviews and publicly available benchmarks for these models remain limited. In this survey, we provide a comprehensive review of multimodal spatial reasoning tasks with large models, categorizing recent progress in multimodal large language models (MLLMs) and introducing open benchmarks for evaluation. We begin by outlining general spatial reasoning, focusing on post-training techniques, explainability, and architecture. Beyond classical 2D tasks, we examine spatial relationship reasoning, scene and layout understanding, as well as visual question answering and grounding in 3D space. We also review advances in embodied AI, including vision-language navigation and action models. Additionally, we consider emerging modalities such as audio and egocentric video, which contribute to novel spatial understanding through new sensors. We believe this survey establishes a solid foundation and offers insights into the growing field of multimodal spatial reasoning. Updated information about this survey, codes and implementation of the open benchmarks can be found at https://github.com/zhengxuJosh/Awesome-Spatial-Reasoning.

多模态空间推理大模型具身AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。