arXiv:2512.15557cs.RO2025-12中稿 · IEEE RA-L被引 3

用视觉语言模型让机器人在陌生环境里更准定位

OMCL: Open-vocabulary Monte Carlo Localization

  • 引入视觉语言特征,实现跨模态观测与地图匹配
  • 支持自然语言初始化全局定位,无需预先标注
  • 在室内(Matterport3D/Replica)和室外(SemanticKITTI)均表现稳健

鲁棒的机器人定位是导航的重要前提,但在地图与传感器数据来自不同模态时变得困难。以往方法通常依赖特定环境、封闭词汇语义或微调特征。本文提出开放词汇蒙特卡洛定位(OMCL),将视觉语言特征融入蒙特卡洛定位框架,可基于相机位姿和由带位姿的RGB-D图像或对齐点云构建的3D地图,准确计算视觉观测的似然。该方法能关联异构模态的观测与地图元素,并通过附近物体的自然语言描述原生初始化全局定位。我们在Matterport3D和Replica上评估了室内场景,在SemanticKITTI上验证了室外场景的泛化能力。

原文摘要 · Abstract (English)

Robust robot localization is an important prerequisite for navigation, but it becomes challenging when the map and robot measurements are obtained from different sensors. Prior methods are often tailored to specific environments, relying on closed-set semantics or fine-tuned features. In this work, we extend Monte Carlo Localization with vision-language features, allowing OMCL to robustly compute the likelihood of visual observations given a camera pose and a 3D map created from posed RGB-D images or aligned point clouds. These open-vocabulary features enable us to associate observations and map elements from different modalities, and to natively initialize global localization through natural language descriptions of nearby objects. We evaluate our approach using Matterport3D and Replica for indoor scenes and demonstrate generalization on SemanticKITTI for outdoor scenes.

机器人定位视觉语言开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。