梳理多模态大模型空间推理能力,揭示当前短板与突破方向
Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- 从认知角度构建空间智能分类体系,按推理复杂度划分任务
- 整合文本、视觉、具身场景三类基准,发现模型表现远低于人类水平
- 对比训练与推理两类提升方法,指出二者互补性,适合新研究者参考
空间推理要求感知和操作三维世界中的空间关系,是人类智能的基础,却仍是多模态大语言模型(MLLMs)的持续挑战。现有综述常按输入模态(如文本、图像、视频或3D)分类进展,但我们认为空间能力不单由输入格式决定。本文提出一种基于认知维度的分类体系,按推理复杂度划分任务,并关联至多种认知功能。将现有基准(涵盖纯文本、视觉语言、具身设置)映射到该体系,回顾评估指标与方法。这一认知视角使跨任务比较更系统,揭示当前模型能力与人类推理间的显著差距。此外,分析了提升空间能力的方法,包括训练型与推理型两类。双重视角厘清了各自优势,发现互补机制。通过梳理任务、基准与最新进展,旨在为新研究者提供全面理解及未来研究的可行动方向。
原文摘要 · Abstract (English)
Spatial reasoning, which requires ability to perceive and manipulate spatial relationships in the 3D world, is a fundamental aspect of human intelligence, yet remains a persistent challenge for Multimodal large language models (MLLMs). While existing surveys often categorize recent progress based on input modality (e.g., text, image, video, or 3D), we argue that spatial ability is not solely determined by the input format. Instead, our survey introduces a taxonomy that organizes spatial intelligence from cognitive aspect and divides tasks in terms of reasoning complexity, linking them to several cognitive functions. We map existing benchmarks across text only, vision language, and embodied settings onto this taxonomy, and review evaluation metrics and methodologies for assessing spatial reasoning ability. This cognitive perspective enables more principled cross-task comparisons and reveals critical gaps between current model capabilities and human-like reasoning. In addition, we analyze methods for improving spatial ability, spanning both training-based and reasoning-based approaches. This dual perspective analysis clarifies their respective strengths, uncovers complementary mechanisms. By surveying tasks, benchmarks, and recent advances, we aim to provide new researchers with a comprehensive understanding of the field and actionable directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。