让大模型看清三维空间,通过几何对齐预训练激活空间感知能力
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

- 设计联合任务强制模型预测稀疏点云与语义标签,唤醒几何意识
- 多层级渐进融合模块提升几何特征利用效率,性能在多个3D任务上显著提升
- 适合关注3D视觉理解、空间推理的开发者与研究者
多模态大语言模型(MLLMs)在语义推理方面表现卓越,但在仅依赖纯RGB输入时难以实现3D空间感知。尽管利用了3D重建模型的隐式几何先验,基于图像的方法仍显著落后于使用显式3D数据的方法。我们认为该差距并非源于几何先验不足,而是训练范式错位:以文本为主导的微调无法激活MLLM中的几何表征。现有方法通常采用简单特征拼接并直接优化下游任务,缺乏几何特异性监督,导致结构信息利用不充分。为此,我们提出GAP-MLLM,一种几何对齐预训练范式,在下游适配前显式激活结构感知。具体地,引入视觉提示联合任务,迫使模型同时预测稀疏点图与语义标签,强化几何意识;并设计多层级渐进融合模块与标记级门控机制,实现几何先验的自适应融合,且不压制语义推理。大量实验表明,GAP-MLLM显著提升几何特征融合效果,在3D视觉定位、3D密集描述和3D视频目标检测任务中均实现稳定性能提升。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models, image-based methods still exhibit a notable performance gap compared to methods using explicit 3D data. We argue that this gap does not arise from insufficient geometric priors, but from a misalignment in the training paradigm: text-dominated fine-tuning fails to activate geometric representations within MLLMs. Existing approaches typically resort to naive feature concatenation and optimize directly for downstream tasks without geometry-specific supervision, leading to suboptimal structural utilization. To address this limitation, we propose GAP-MLLM, a Geometry-Aligned Pre-training paradigm that explicitly activates structural perception before downstream adaptation. Specifically, we introduce a visual-prompted joint task that compels the MLLMs to predict sparse pointmaps alongside semantic labels, thereby enforcing geometric awareness. Furthermore, we design a multi-level progressive fusion module with a token-level gating mechanism, enabling adaptive integration of geometric priors without suppressing semantic reasoning. Extensive experiments demonstrate that GAP-MLLM significantly enhances geometric feature fusion and consistently enhances performance across 3D visual grounding, 3D dense captioning, and 3D video object detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。