通过隐式结构引导提升多模态大模型的3D空间推理能力
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- 利用3D重建的结构感知实现隐式空间引导,无需依赖点云渲染
- 在ScanRefer、Nr3D、Sr3D上显著超越现有方法,兼顾性能与效率
- 适合需要高效3D视觉定位的机器人和具身智能研究者
3D视觉定位(3DVG)旨在根据自然语言描述在3D场景中定位物体,是具身智能与机器人领域的基础任务。近年来,多模态大模型(MLLMs)的发展推动了其向3DVG的拓展。然而,现有MLLM主要处理2D视觉输入,仅凭有限视角难以理解场景的3D空间结构。当前方法多依赖重建点云的视角相关渲染提供显式结构引导,存在效率低、空间推理受限的问题。为此,我们提出S²-MLLM,一种通过隐式空间推理增强MLLM空间理解能力的高效框架。该方法利用前馈3D重建中的结构感知,在训练中获取3D结构理解,从而无需依赖低效的点云重构进行推理。此外,我们设计结构增强模块(SE),结合视图内与视图间注意力机制,捕捉单视图内部及多视图间的关联,并融合多层次位置编码,将视觉表征与空间位置和视角信息对齐,实现更精准的结构理解。大量实验表明,S²-MLLM在扫描参照(ScanRefer)、Nr3D与Sr3D数据集上均取得显著优于现有方法的性能,兼具优越性、泛化性和效率。代码将在论文接受后公开。
原文摘要 · Abstract (English)
3D Visual Grounding (3DVG) focuses on locating objects in 3D scenes based on natural language descriptions, serving as a fundamental task for embodied AI and robotics. Recent advances in Multi-modal Large Language Models (MLLMs) have motivated research into extending them to 3DVG. However, MLLMs primarily process 2D visual inputs and struggle with understanding 3D spatial structure of scenes solely from these limited perspectives. Existing methods mainly utilize viewpoint-dependent rendering of reconstructed point clouds to provide explicit structural guidance for MLLMs in 3DVG tasks, leading to inefficiency and limited spatial reasoning. To address this issue, we propose S$^2$-MLLM, an efficient framework that enhances spatial reasoning in MLLMs through implicit spatial reasoning. We introduce a spatial guidance strategy that leverages the structure awareness of feed-forward 3D reconstruction. By acquiring 3D structural understanding during training, our model can implicitly reason about 3D scenes without relying on inefficient point cloud reconstruction. Moreover, we propose a structure-enhanced module (SE), which first employs intra-view and inter-view attention mechanisms to capture dependencies within views and correspondences across views. The module further integrates multi-level position encoding to associate visual representations with spatial positions and viewpoint information, enabling more accurate structural understanding. Extensive experiments demonstrate that S$^2$-MLLM unifies superior performance, generalization, and efficiency, achieving significant performance over existing methods across the ScanRefer, Nr3D, and Sr3D datasets. Code will be available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。