用物理感知增强视觉语言模型,让机器人更安全地穿越复杂户外地形。
Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments
- 结合本体感知数据动态校准视觉语言模型的语义理解。
- 在真实户外环境中实现最高50%的导航成功率提升。
- 适合需要应对泥地、雪地等复杂地形的移动机器人研发者。
我们提出一种新型自主机器人导航算法VLM-GroNav,用于处理户外多变的地形可通行性问题。该方法利用视觉语言模型(VLMs)并引入物理接地机制,评估地形的变形性和滑溜性等内在属性。通过本体感知传感直接获取这些物理特性,增强对地形的语义理解。采用上下文学习将本体数据与VLM的语义理解融合,实现基于机器人实时物理交互的可通行性动态更新。更新后的估计结果用于指导局部和全局规划器进行实时轨迹重规划。我们在腿式机器人(Ghost Vision 60)和轮式机器人(Clearpath Husky)上验证了该方法,在多种真实户外场景中测试了不同柔性和滑溜地形的表现。实际应用中,相较于现有最先进方法,导航成功率最高提升达50%。
原文摘要 · Abstract (English)
We present a novel autonomous robot navigation algorithm for outdoor environments that is capable of handling diverse terrain traversability conditions. Our approach, VLM-GroNav, uses vision-language models (VLMs) and integrates them with physical grounding that is used to assess intrinsic terrain properties such as deformability and slipperiness. We use proprioceptive-based sensing, which provides direct measurements of these physical properties, and enhances the overall semantic understanding of the terrains. Our formulation uses in-context learning to ground the VLM's semantic understanding with proprioceptive data to allow dynamic updates of traversability estimates based on the robot's real-time physical interactions with the environment. We use the updated traversability estimations to inform both the local and global planners for real-time trajectory replanning. We validate our method on a legged robot (Ghost Vision 60) and a wheeled robot (Clearpath Husky), in diverse real-world outdoor environments with different deformable and slippery terrains. In practice, we observe significant improvements over state-of-the-art methods by up to 50% increase in navigation success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。