arXiv:2605.25334cs.CV2026-05

仅用RGB图像实现三维结构与尺度感知,提升空间理解能力。

Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence

论文配图:Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence
图 1 · 摘自论文原文
  • 双路径架构分离提取度量尺度与结构线索,避免信息干扰。
  • 在7个空间智能基准上达到当前最优性能,优于现有模型。
  • 适合需要高效空间推理的机器人、自动驾驶等应用。

从2D视觉输入中理解物理世界依赖两种互补的几何知识:整体3D结构感知和精细度量尺度估计。现有多模态大语言模型通常只关注其中一方面,需额外输入深度图或点云,带来高计算开销并继承上游预测模型的泛化局限。我们提出GAMSI,一种仅以RGB图像为输入的双路径几何感知多模态大语言模型,通过统一自回归主干内化两类几何先验。具体地,引入度量-结构解耦查询(MSDQ),利用两组可学习查询分别从共享视觉上下文中提取密集度量信号与稀疏结构线索,并通过任务解耦注意力掩码防止路径间污染。在此基础上,专家引导视觉定位(EVG)模块将聚合线索投影回帧级视觉特征,与视觉基础模型对齐,后者仅作为训练阶段监督,而非模型输入。我们还构建了一个包含152,776样本、涵盖13类任务和三种视觉模态的多任务空间指令微调数据集(MTS),整合自六个公开数据集。采用两阶段课程训练策略,GAMSI在七个空间智能基准上取得当前最优表现。

原文摘要 · Abstract (English)

Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metric scale estimation. Existing multimodal large language models (MLLMs) typically address only one facet, ingesting either depth maps or point clouds as additional model inputs, which incurs substantial computational overhead and inherits the generalization limitations of upstream prediction models. We propose GAMSI, a dual-pathway Geometry-Aware MLLM for Spatial Intelligence that takes only RGB images as input while internalizing both forms of geometric prior within a unified autoregressive backbone. Specifically, we introduce Metric-Structure Decoupled Queries (MSDQ) which employ two groups of learnable queries to respectively extract dense metric signals and sparse structural cues from the shared visual context, with a task-decoupled attention mask further preventing the two pathways from contaminating each other. Building on this, an Expert-Guided Visual Grounding (EVG) module projects the aggregated cues back to frame-level visual features and aligns them with vision foundation models, which serve purely as training-time supervision, rather than as model inputs. We further build a multi-task spatial instruction-tuning dataset (MTS) comprising 152{,}776 samples spanning 13 task types and three visual modalities, consolidated from six public datasets. Trained with a two-stage curriculum, GAMSI achieves state-of-the-art performance on seven spatial intelligence benchmarks.

空间理解多模态模型几何先验视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。