让多模态大模型自主生成3D空间表示,实现可解释的三维感知。
SpatialSV: Internalizing Interpretable 3D Spatial Awareness in MLLMs via Task-Oriented Visual Supervision

- 通过任务导向的视觉监督,让模型主动将2D特征转为深度图、相机位姿等3D表示
- 在多个基准上显著提升模型空间理解能力,且3D重建结果可直观诊断模型表现
- 适用于需要可解释性的场景,如自动驾驶、机器人导航等
解锁多模态大语言模型(MLLMs)的空间智能对于理解与交互三维世界至关重要。现有方法通常依赖外部工具注入空间先验,带来显著推理开销;或通过隐式特征蒸馏,缺乏可解释性与细粒度几何约束。为此,我们提出SpatialSV框架,旨在内化鲁棒的3D空间感知能力,同时保持内在可解释性。不同于被动特征模仿,SpatialSV采用任务导向的视觉监督,迫使模型主动将2D视觉特征提升为显式的3D表示,包括深度图、相机位姿和点云。关键的是,这一2D到3D的转换过程为模型表征提供了透明窗口:生成的3D重建结果可作为直观代理,用于可视化与诊断模型内在空间知识的质量。在多个模型与基准上的大量实验表明,SpatialSV能有效增强并解释MLLMs的空间智能。此外,该框架在半监督设置中展现出强泛化能力,验证了其利用未标注视觉数据实现可扩展、可解释的空间表征学习的潜力。
原文摘要 · Abstract (English)
Unlocking the spatial intelligence of multimodal large language model (MLLMs) is crucial for understanding and interacting with the 3D world. Prevailing approaches typically inject spatial priors via external tools, which impose significant inference overhead, or rely on latent feature distillation, which remains uninterpretable and lacks fine-grained geometric constraints. To address these issues, we propose SpatialSV, a framework designed to internalize robust 3D spatial awareness within MLLMs while simultaneously offering inherent interpretability. Deviating from passive feature imitation, SpatialSV employs task-oriented visual supervision, compelling the model to actively lift its 2D visual features into explicit 3D representations, including depth maps, camera poses, and point clouds. Crucially, this 2D-to-3D lifting process provides a transparent window into the model's representations: the resulting 3D reconstructions serve as an intuitive proxy for visualizing and diagnosing the quality of the model's intrinsic spatial knowledge. Extensive experiments across multiple models and benchmarks demonstrate the effectiveness of SpatialSV in enhancing and interpreting MLLMs' spatial intelligence. Furthermore, the framework exhibits strong generalization in semi-supervised settings, validating its potential to leverage unlabeled visual data for scalable, interpretable spatial representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。