用语义统计替代固定地标,实现低成本硬件下的实时精准定位
ShelfAware: Real-Time Semantic Localization in Quasi-Static Environments with Low-Cost Sensors
- 将场景语义建模为类别分布,通过逆向语义推测生成定位假设
- 在模拟零售环境成功率97%,动态遮挡下仍保持66%追踪成功率
- 无需基础设施,适合移动机器人和助人设备在动态环境中使用
许多室内工作空间具有准静态特性:全局几何结构稳定,但局部语义持续变化,导致重复几何、动态杂物和感知噪声,使传统视觉定位失效。本文提出ShelfAware,一种基于语义粒子滤波的鲁棒全局定位方法,将场景语义视为物体类别上的统计证据,而非固定数量的地标。ShelfAware融合深度似然与以类别为中心的语义相似性,并利用预构建的语义视角库,在蒙特卡洛定位(MCL)中进行逆向语义提议,实现在低成本纯视觉硬件上的快速、精准假设生成。为验证感知无关的可扩展性,我们在两个场景中评估该方法:在严格控制的模拟零售环境,全局定位成功率达97%,并在购物车、穿戴式设备及动态遮挡条件下维持最高66%的追踪成功率;在3,500平方英尺的真实运营超市中,借助开放词汇视觉管道,ShelfAware显著优于几何与固定数量语义基线。通过分布式建模语义并利用逆向提议,ShelfAware有效解决几何混淆问题,为移动与辅助机器人在动态现实环境中提供无基础设施的通用定位基础。
原文摘要 · Abstract (English)
Many indoor workspaces are quasi-static: their global geometric layout is stable, but local semantics change continually, producing repetitive geometry, dynamic clutter, and perceptual noise that defeat standard vision-based localization. We present ShelfAware, a semantic particle filter for robust global localization that treats scene semantics as statistical evidence over object categories rather than fixed quantity landmarks. ShelfAware fuses a depth likelihood with a category-centric semantic similarity and uses a precomputed bank of semantic viewpoints to perform inverse semantic proposals inside Monte Carlo Localization (MCL), yielding fast, targeted hypothesis generation on low-cost, vision-only hardware. To demonstrate perception-agnostic scalability, we evaluate ShelfAware across two domains. In a rigorously controlled mock retail environment, ShelfAware achieves a 97% global localization success rate, maintaining the highest tracking success (66%) across cart, wearable, and dynamic occlusion conditions. Furthermore, in a 3,500 sq. ft. operational grocery store leveraging an open-vocabulary vision pipeline, ShelfAware significantly outperforms both geometric and fixed-quantity semantic baselines. By modeling semantics distributionally and leveraging inverse proposals, ShelfAware resolves geometric aliasing, providing an infrastructure-free building block for mobile and assistive robots in dynamic real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。