arXiv:2605.30506cs.ROcs.CV2026-05

用视觉语言模型提升机器人在杂乱环境中的定位精度

VLM-GLoc: Vision-Language Model Enhanced Monte Carlo Localization for Robust Semantic Global Localization in Cluttered Quasi-Static Environments

论文配图:VLM-GLoc: Vision-Language Model Enhanced Monte Carlo Localization for Robust Semantic Global Localization in Cluttered Quasi-Static Environments
图 1 · 摘自论文原文
  • 用开放词汇的视觉语言模型做语义感知,统一处理复杂场景
  • 在超市和实验室分别达到70%和74%定位成功率,显著优于传统方法
  • 适合需要在重复布局中精准定位的移动机器人应用

在几何混淆、准静态的室内环境中(如超市、办公室、实验室和医院),移动机器人实现鲁棒的全局定位面临巨大挑战。超市中平行货架与长尾产品分布,以及办公/实验室内重复的桌椅、显示器、门等家具,导致几何与语义双重模糊。传统方法依赖特定几何特征或领域专用视觉管道,难以应对长尾语义分布和临时视觉干扰。本文提出VLM-GLoc,一种基于开放词汇视觉语言模型(VLM)的分层语义蒙特卡洛定位方法。通过文本到地图检索生成初始粒子,利用VLM提取高区分性语义特征、隐式过滤模糊或动态物体,并实现目标数据增强的持久性推理。在两个真实环境(3,500平方英尺超市与3,700平方英尺实验室)及两种平台(手机与四足机器人)上评估,定位成功率分别达70%和74%,显著超越仅依赖几何信息或领域特定基线的方法。

原文摘要 · Abstract (English)

Global localization in geometrically aliased, quasi-static environments such as grocery stores, offices, schools, and hospitals poses a significant challenge for mobile robots. Grocery stores with parallel aisles and a long tailed distribution of products, as well as offices and labs with repetitive furniture such as chairs, desks, monitors, and doors, exemplify common indoor environments that present geometric and even semantic ambiguity. Traditional approaches rely either on distinct geometric features or on domain-specific vision pipelines that struggle with long-tail semantic distributions and transient visual clutter. We present VLM-GLoc, a method for hierarchical semantic Monte Carlo Localization (MCL) that leverages open-vocabulary Vision-Language Models (VLMs) as a unified semantic observation front-end. We hypothesize a three-fold benefit from VLMs: (1) extracting highly discriminative rich text features, (2) implicit quality filtering of blurry or dynamic objects, and (3) permanence reasoning for targeted data augmentation. We introduce an inverse semantic proposal mechanism that seeds particles via text-to-map retrieval. Evaluated across two real-world environments with different characteristics and two different platforms: a 3,500 sq. ft. grocery store with a cellphone and a 3,700 sq. ft. lab space with a quadruped, VLM-GLoc achieves 70% and 74% global localization success respectively, substantially outperforming traditional geometry-only and domain-specific baselines.

语义定位视觉语言模型机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。