arXiv:2410.03878cs.CV2024-10ICLR被引 27

构建3D场景的沉浸式理解数据集,提升大模型空间推理能力

SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models

  • 构建沉浸式3D数据集Spartun3D,含多种空间推理任务
  • 提出新对齐模块,增强视觉与语言的空间一致性
  • 适合做3D场景理解、空间推理的AI研究者参考

将三维世界融入大语言模型(3D-based LLMs)是实现三维场景理解的前沿方向。然而,现有3D-based LLMs在情境化空间理解方面存在两大不足:一是现有3D数据集从全局视角构建,缺乏情境上下文;二是模型架构中三维场景的空间表示与自然语言之间缺乏显式对齐,限制了精准空间推理任务的表现。为此,我们提出可扩展的情境化3D数据集Spartun3D,涵盖多种情境空间推理任务。同时,基于现有3D-based LLM,设计Spartun3D-LLM,引入新颖的情境空间对齐模块,旨在强化三维视觉表征与对应文本描述之间的对齐。实验表明,所提出的数据集和对齐模块均显著提升了3D-based LLM的情境空间理解能力。

原文摘要 · Abstract (English)

Integrating the 3D world into large language models (3D-based LLMs) has been a promising research direction for 3D scene understanding. However, current 3D-based LLMs fall short in situated understanding due to two key limitations: 1) existing 3D datasets are constructed from a global perspective of the 3D scenes and lack situated context. 2) the architectures of existing 3D-based LLMs lack explicit alignment between the spatial representations of 3D scenes and natural language, limiting their performance in tasks requiring precise spatial reasoning. We address these issues by introducing a scalable situated 3D dataset, named Spartun3D, that incorporates various situated spatial reasoning tasks. Furthermore, we propose Spartun3D-LLM, built on an existing 3D-based LLM but integrated with a novel situated spatial alignment module, aiming to enhance the alignment between 3D visual representations and their corresponding textual descriptions. Experimental results demonstrate that both our proposed dataset and alignment module significantly enhance the situated spatial understanding of 3D-based LLMs.

3D理解空间推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。