arXiv:2506.16623cs.ROcs.AI2025-06被引 5

用历史上下文增强视觉语言模型,让机器人更聪明地找物导航

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

  • 给VLM加入动作历史,动态生成导航指导分数
  • 在HM3D上达46%成功率,24.8%路径加权成功率
  • 适合做零样本机器人导航的算法研究者参考

物体目标导航(ObjectNav)要求机器人在未见过的环境中寻找物体,需要复杂的推理能力。尽管视觉语言模型(VLMs)展现出潜力,但现有方法通常仅用其视觉-语言嵌入进行物体-场景相似性匹配,未能深入利用推理能力,导致上下文理解不足,并出现重复导航等问题。本文提出一种新的零样本物体导航框架,首次引入动态、历史感知的提示机制,将VLM推理深度融入基于前沿的探索过程。核心创新在于向VLM提供动作历史信息,使其能生成导航动作的语义引导评分,主动避免决策循环。同时引入由VLM辅助的路径点生成机制,优化对目标物体的最终接近。在Habitat平台的HM3D数据集上评估,该方法取得46%的成功率(SR)和24.8%的路径长度加权成功率(SPL),结果与当前最先进零样本方法相当,验证了历史增强型VLM提示策略在提升机器人导航鲁棒性与上下文感知能力方面的显著潜力。

原文摘要 · Abstract (English)

Object Goal Navigation (ObjectNav) challenges robots to find objects in unseen environments, demanding sophisticated reasoning. While Vision-Language Models (VLMs) show potential, current ObjectNav methods often employ them superficially, primarily using vision-language embeddings for object-scene similarity checks rather than leveraging deeper reasoning. This limits contextual understanding and leads to practical issues like repetitive navigation behaviors. This paper introduces a novel zero-shot ObjectNav framework that pioneers the use of dynamic, history-aware prompting to more deeply integrate VLM reasoning into frontier-based exploration. Our core innovation lies in providing the VLM with action history context, enabling it to generate semantic guidance scores for navigation actions while actively avoiding decision loops. We also introduce a VLM-assisted waypoint generation mechanism for refining the final approach to detected objects. Evaluated on the HM3D dataset within Habitat, our approach achieves a 46% Success Rate (SR) and 24.8% Success weighted by Path Length (SPL). These results are comparable to state-of-the-art zero-shot methods, demonstrating the significant potential of our history-augmented VLM prompting strategy for more robust and context-aware robotic navigation.

机器人导航视觉语言模型零样本历史感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。