用视觉语言模型提升多场景相机定位的泛化能力
MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization
- 利用预训练视觉语言模型引入世界知识
- 在7Scenes和Cambridge Landmarks上实现领先精度
- 支持室内外场景,适合需要跨环境定位的应用
相机重定位是现代计算机视觉的核心能力,可从图像中准确确定相机的位置与姿态(6-DoF),对增强现实(AR)、混合现实(MR)、自动驾驶、配送无人机及机器人导航至关重要。与传统仅针对单场景回归姿态的深度学习方法不同,MVL-Loc提出一种端到端的多场景6-DoF相机重定位框架,借助预训练视觉语言模型(VLMs)中的世界知识,并融合多模态数据,实现对室内外复杂场景的强泛化能力。自然语言被用作引导指令,促进对复杂场景的语义理解与物体间空间关系建模。在7Scenes和Cambridge Landmarks数据集上的大量实验表明,MVL-Loc在真实多场景环境下具备鲁棒性,且在位置与姿态估计上均达到当前最优性能。
原文摘要 · Abstract (English)
Camera relocalization, a cornerstone capability of modern computer vision, accurately determines a camera's position and orientation (6-DoF) from images and is essential for applications in augmented reality (AR), mixed reality (MR), autonomous driving, delivery drones, and robotic navigation. Unlike traditional deep learning-based methods that regress camera pose from images in a single scene, which often lack generalization and robustness in diverse environments, we propose MVL-Loc, a novel end-to-end multi-scene 6-DoF camera relocalization framework. MVL-Loc leverages pretrained world knowledge from vision-language models (VLMs) and incorporates multimodal data to generalize across both indoor and outdoor settings. Furthermore, natural language is employed as a directive tool to guide the multi-scene learning process, facilitating semantic understanding of complex scenes and capturing spatial relationships among objects. Extensive experiments on the 7Scenes and Cambridge Landmarks datasets demonstrate MVL-Loc's robustness and state-of-the-art performance in real-world multi-scene camera relocalization, with improved accuracy in both positional and orientational estimates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。