让机器人导航时自动识别语言提示的不确定性,更聪明地找东西。
Uncertainty-Informed Active Perception for Open Vocabulary Object Goal Navigation
- 用概率模型量化视觉语言模型的语义不确定度
- 在地图中融合不确定度信息,提升空间理解能力
- 无需复杂提示工程,适合真实场景的智能导航
移动机器人在室内环境中越来越依赖视觉-语言模型来识别图像中的高层语义线索,如物体类别。这类模型为对象目标导航(ObjectNav)等任务带来显著进步,即机器人需通过探索环境定位自然语言描述的物体。当前方法严重依赖提示工程进行感知,且未解决因提示表述差异引发的语义不确定性问题。忽略语义不确定性会导致次优探索行为,从而限制性能表现。为此,我们提出一种面向室内环境对象目标导航的语义不确定性感知主动感知框架。引入新型概率传感器模型,用于量化视觉-语言模型中的语义不确定性,并将其融入概率性几何-语义地图以增强空间理解。基于该地图,设计了一种基于不确定性感知多臂赌博机目标的前沿探索规划器,实现高效对象搜索。实验结果表明,本方法在无需大量提示工程的情况下,达到了与现有顶尖方法相当的对象导航成功率。
原文摘要 · Abstract (English)
Mobile robots exploring indoor environments increasingly rely on vision-language models to perceive high-level semantic cues in camera images, such as object categories. Such models offer the potential to substantially advance robot behaviour for tasks such as object-goal navigation (ObjectNav), where the robot must locate objects specified in natural language by exploring the environment. Current ObjectNav methods heavily depend on prompt engineering for perception and do not address the semantic uncertainty induced by variations in prompt phrasing. Ignoring semantic uncertainty can lead to suboptimal exploration, which in turn limits performance. Hence, we propose a semantic uncertainty-informed active perception pipeline for ObjectNav in indoor environments. We introduce a novel probabilistic sensor model for quantifying semantic uncertainty in vision-language models and incorporate it into a probabilistic geometric-semantic map to enhance spatial understanding. Based on this map, we develop a frontier exploration planner with an uncertainty-informed multi-armed bandit objective to guide efficient object search. Experimental results demonstrate that our method achieves ObjectNav success rates comparable to those of state-of-the-art approaches, without requiring extensive prompt engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。