用几何标记3D场景,让多模态大模型更好理解空间关系。
Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations

- 给图像中物体加唯一标识并编码3D几何属性为文本引用。
- 零样本下在VSI-Bench和MindCube上分别提升9%和12%性能。
- 适合需要强空间推理能力的视觉问答与机器人导航任务。
尽管多模态大语言模型(MLLMs)在二维视觉理解上已取得显著进展,但其对三维空间的推理能力仍有限。为弥补这一差距,我们提出几何参考3D场景表示(GR3D)。给定一组输入图像,GR3D为图像中的物体分配唯一ID,并将其3D几何属性编码为以这些ID索引的文本引用。该表示使MLLM能利用其强大的语言数学推理能力解析3D线索,同时紧密耦合地分析2D视觉特征。我们提出一种基于GR3D的简单有效方法,无需额外训练,可直接适配多种MLLM。在零样本设置下,该方法在挑战性空间推理基准测试中表现显著提升:在VSI-Bench上使GPT-5性能提升9%,在MindCube上提升12%。定性研究进一步表明,GR3D使MLLM能在极稀疏视角输入下完成复杂空间推理。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space remains limited. To address this gap, we introduce geometrically referenced 3D scene representations (GR3D). Given a set of input images, GR3D annotates objects in the images with unique IDs and encodes their 3D geometric attributes as textual references indexed by these IDs. This representation enables MLLMs to interpret 3D cues using their advanced language-based skills in mathematical reasoning, while concurrently analyzing 2D visual features in a tightly coupled way. We present a simple yet effective approach based on GR3D, which requires no additional training and is readily applicable to different MLLMs. Implemented in a zero-shot setting, our approach yields substantial improvements on challenging spatial reasoning benchmarks, boosting GPT-5 performance by 9% on VSI-Bench and 12% on MindCube. Qualitative studies further demonstrate that GR3D empowers MLLMs to perform complex spatial reasoning with highly sparse input views.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。