用2D视频训练模型,让大语言模型学会空间感知。
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

- 从2D视频中蒸馏3D几何知识,重塑模型内部表示
- 在多个空间推理任务上达到当前最优性能
- 适合需要空间理解的视觉语言模型研究者
多模态大语言模型(MLLM)擅长2D语义理解,但缺乏内在3D意识,导致视频帧间几何与空间一致性缺失。由于大规模3D数据稀缺,我们提出GeoVR框架,仅使用2D视频序列学习几何表征。该方法通过蒸馏预训练3D基础模型的几何知识,重构MLLM的语义潜在空间,以激发空间智能。采用多目标学习策略,基于四个互补的几何目标:(1) 估计帧间相机位姿以嵌入视角变化动态,(2) 回归密集深度图以锚定物理距离,(3) 预测度量尺度因子实现真实世界校准,(4) 蒸馏多尺度3D特征以对齐中间特征空间。在这些显式物理与几何约束下,模型内部表示自然发展出强3D感知能力。大量实验表明,GeoVR在空间推理基准测试中达到最先进水平,确立了赋予基础模型空间智能的新范式。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to maintain geometric and spatial consistency across video frames. Given the scarcity of large-scale 3D data, we present GeoVR, a novel framework that learns geometric representations using purely 2D video sequences. This approach effectively restructures the semantic latent space within MLLMs to unlock spatial intelligence. Rather than employing superficial feature mixing, GeoVR reshapes the internal representations of the MLLM by distilling geometry knowledge from pre-trained 3D foundation models. This is accomplished through a multi-objective learning strategy driven by four complementary geometric targets: (1) estimating inter-frame camera poses to embed varying viewpoint dynamics, (2) regressing dense depth maps to anchor physical distances, (3) predicting a metric scale factor for real-world calibration, and (4) distilling multi-scale 3D features to align the intermediate feature space. Guided by these explicit physical and geometric constraints, the model's internal representations naturally develop strong 3D awareness. Extensive experiments on spatial reasoning benchmarks demonstrate that GeoVR achieves state-of-the-art performance, establishing a new paradigm for endowing foundation models with spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。