让大模型学会三维空间推理,超越GPT-4o性能8.7%
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
- 构建两类3D感知数据集,支持物体位置与空间关系理解
- 在真实图像上首次引入含3D朝向关系的视觉问答数据
- 系统优化架构与训练策略,实现优于GPT-4o的3D推理能力
人类天然具备三维空间理解能力,可进行车辆交叉碰撞等复杂推理。现有大型多模态模型(LMMs)缺乏此类3D空间推理能力,主要源于3D训练数据稀缺及模型设计对2D数据的偏向。本文系统研究3D感知数据、架构与训练方案的影响,提出SpatialLLM——一个具备先进3D空间推理能力的大规模多模态模型。为解决数据瓶颈,我们构建两类3D感知训练数据:(1) 侧重物体3D位置与朝向的探针数据;(2) 支持复杂空间关系对话的交互数据。尤为关键的是,我们首次在真实图像上构建包含3D朝向关系的视觉问答数据。通过系统融合数据、架构与训练设计,提供实现卓越3D推理能力的优化路径。实验表明,SpatialLLM在3D推理任务上超越GPT-4o达8.7%。本研究的实证设计与发现为后续研究提供重要参考。
原文摘要 · Abstract (English)
Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial reasoning. This limitation stems from the scarcity of 3D training data and the bias in current model designs toward 2D data. In this paper, we systematically study the impact of 3D-informed data, architecture, and training setups, introducing SpatialLLM, a large multi-modal model with advanced 3D spatial reasoning abilities. To address data limitations, we develop two types of 3D-informed training datasets: (1) 3D-informed probing data focused on object's 3D location and orientation, and (2) 3D-informed conversation data for complex spatial relationships. Notably, we are the first to curate VQA data that incorporate 3D orientation relationships on real images. Furthermore, we systematically integrate these two types of training data with the architectural and training designs of LMMs, providing a roadmap for optimal design aimed at achieving superior 3D reasoning capabilities. Our SpatialLLM advances machines toward highly capable 3D-informed reasoning, surpassing GPT-4o performance by 8.7%. Our systematic empirical design and the resulting findings offer valuable insights for future research in this direction. Our project page is available at: https://3d-spatial-reasoning.github.io/spatial-llm/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。