让视觉模型理解相机参数,提升跨视角空间推理能力
On the Generalization Capacities of MLLMs for Spatial Intelligence
- 通过注入相机内参嵌入,让每个视觉令牌感知拍摄视角
- 在跨相机测试中,性能显著优于传统方法,泛化能力提升明显
- 适合研究多视角视觉理解与具身智能的学者
直接处理RGB输入的多模态大语言模型(MLLMs)在3D定位与导航任务中展现出巨大潜力。然而,我们指出仅依赖RGB输入的方法在跨摄像头泛化上存在根本缺陷:忽略相机参数会将物体物理属性与视角混淆,导致模型过拟合训练时的相机分布,而非学习通用的3D几何规律。为此,我们提出相机感知型MLLM框架,通过三项机制实现可泛化的空间推理:(i) 通过密集嵌入注入相机内参,使每个视觉令牌受视角条件约束;(ii) 引入相机感知数据增强,合成不同相机参数,迫使模型剥离视角信息;(iii) 从3D视觉基础模型中蒸馏几何先验。大量实验表明,相机感知的MLLM在空间定位任务的跨摄像头泛化测试中显著优于基准模型,证明相机感知不仅是优势,更是构建鲁棒、通用空间智能的必要条件。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) that directly process RGB inputs for tasks like 3D localization and navigation have shown remarkable potential. However, we argue that these RGB-only approaches are fundamentally flawed in their ability to generalize across cameras. By ignoring camera parameters, they entangle an object's physical properties with the camera's perspective, creating an irresolvable ambiguity. We show this leads MLLMs to overfit to the training camera distribution, rather than learning true and generalizable 3D geometric principles. To address this, we propose Camera-Aware MLLM framework for spatial MLLMs. It learns generalizable spatial reasoning by: (i) injecting camera intrinsics via a dense embedding that conditions each visual token; (ii) introducing a camera-aware data augmentation strategy that synthetically varies camera parameters, forcing the model to disentangle camera properties from scene content; and (iii) distilling geometric priors from a 3D vision foundation model. Extensive experiments demonstrate that camera-aware MLLMs substantially outperform their naive counterparts, particularly in cross-camera generalization tests on spatially-grounded tasks, indicating that camera-awareness is not only beneficial but also a prerequisite for robust and generalizable spatial intelligence in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。