arXiv:2510.22917cs.ROcs.AI2025-10被引 3

融合视觉语言模型,让机器人同时看局部和全局信息导航

HyPerNav: Hybrid Perception for Object-Oriented Navigation in Unknown Environment

  • 用视觉语言模型统一处理摄像头和俯视地图数据
  • 在仿真和真实环境均超越现有方法,定位更准更快
  • 适合做智能机器人导航的开发者参考

目标导向导航(ObjNav)使机器人能在未知环境中直接、自主地寻找到目标物体。有效的感知能力对未知环境中的自主导航至关重要。虽然来自RGB-D传感器的自我中心观测提供丰富的局部信息,实时俯视地图则提供有价值的全局上下文,但大多数现有研究仅关注单一信息源,很少整合这两种互补的感知模态。尽管人类自然会同时关注两者。随着视觉-语言模型(VLMs)的快速发展,我们提出混合感知导航(HyPerNav),利用VLM强大的推理与视觉-语言理解能力,联合感知局部与全局信息,以提升未知环境中导航的有效性和智能化水平。在大规模仿真评估和真实世界验证中,该方法在多个主流基线中表现最优。得益于混合感知策略,我们的方法能捕捉更丰富的线索,更有效地找到目标物体,同时利用自我中心观测与俯视地图的信息理解。消融实验证明,两种感知方式均对导航性能有贡献。

原文摘要 · Abstract (English)

Objective-oriented navigation(ObjNav) enables robot to navigate to target object directly and autonomously in an unknown environment. Effective perception in navigation in unknown environment is critical for autonomous robots. While egocentric observations from RGB-D sensors provide abundant local information, real-time top-down maps offer valuable global context for ObjNav. Nevertheless, the majority of existing studies focus on a single source, seldom integrating these two complementary perceptual modalities, despite the fact that humans naturally attend to both. With the rapid advancement of Vision-Language Models(VLMs), we propose Hybrid Perception Navigation (HyPerNav), leveraging VLMs' strong reasoning and vision-language understanding capabilities to jointly perceive both local and global information to enhance the effectiveness and intelligence of navigation in unknown environments. In both massive simulation evaluation and real-world validation, our methods achieved state-of-the-art performance against popular baselines. Benefiting from hybrid perception approach, our method captures richer cues and finds the objects more effectively, by simultaneously leveraging information understanding from egocentric observations and the top-down map. Our ablation study further proved that either of the hybrid perception contributes to the navigation performance.

机器人导航视觉语言模型多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。