让多模态大模型学会理解3D空间,提升室内场景的几何推理能力。
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
- 构建新数据集CA-VQA,覆盖空间关系、尺寸距离估计等3D任务。
- 训练出的MM-Spatial模型在3D理解任务上达到当前最佳性能。
- 仅用数据即可让模型具备接近专用深度估计模型的感知能力。
多模态大语言模型在2D视觉理解上表现优异,但在3D空间推理方面仍受限。本文利用大规模高质量3D场景数据及开放集标注,提出1)一个新型有监督微调数据集,2)一个聚焦室内场景的新评估基准。所构建的Cubify Anything VQA(CA-VQA)数据集涵盖多种空间任务,包括空间关系预测、度量尺度与距离估计、3D定位等。实验表明,基于CA-VQA训练的MM-Spatial模型不仅具备强大泛化能力,还在多个3D空间理解基准上达到当前最优性能。我们验证了引入度量深度信息与多视角输入(在CA-VQA中提供)对3D理解的提升作用,并证明仅靠数据即可使模型实现与专用单目深度估计模型相当的深度感知能力。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) excel at 2D visual understanding but remain limited in their ability to reason about 3D space. In this work, we leverage large-scale high-quality 3D scene data with open-set annotations to introduce 1) a novel supervised fine-tuning dataset and 2) a new evaluation benchmark, focused on indoor scenes. Our Cubify Anything VQA (CA-VQA) data covers diverse spatial tasks including spatial relationship prediction, metric size and distance estimation, and 3D grounding. We show that CA-VQA enables us to train MM-Spatial, a strong generalist MLLM that also achieves state-of-the-art performance on 3D spatial understanding benchmarks, including our own. We show how incorporating metric depth and multi-view inputs (provided in CA-VQA) can further improve 3D understanding, and demonstrate that data alone allows our model to achieve depth perception capabilities comparable to dedicated monocular depth estimation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。