用视觉语言模型实现行星探测车智能导航,高效又安全。
VLM-Empowered Multi-Mode System for Efficient and Safe Planetary Navigation
- 通过视觉语言模型理解地形复杂度,动态切换导航模式。
- 长距离穿越效率提升79.5%,同时保持对地形危险的规避能力。
- 适合需要自主导航的复杂环境探测任务,如火星或月球巡视。
日益复杂的行星探测环境要求更灵活的巡视车导航策略。本文提出一种由视觉语言模型(VLM)驱动的多模式系统,实现行星巡视车高效且安全的自主导航。该系统利用视觉语言模型解析图像输入,实现对地形复杂度的人类级理解。基于复杂度分类,系统自动切换至最适配的导航模式,包含针对不同地形设计的感知、建图与规划模块,以提前穿越前方地形并抵达下一个航点。通过将局部导航系统与地图服务器及全局航点生成模块集成,巡视车可在复杂场景中完成长距离导航任务。系统在多种仿真环境中评估,相较于单模式保守导航方法,在包含多种障碍物的长距离穿越中,效率提升79.5%,同时保持对地形风险的规避能力,保障巡视车安全。更多系统信息详见 https://chengsn1234.github.io/multi-mode-planetary-navigation/。
原文摘要 · Abstract (English)
The increasingly complex and diverse planetary exploration environment requires more adaptable and flexible rover navigation strategy. In this study, we propose a VLM-empowered multi-mode system to achieve efficient while safe autonomous navigation for planetary rovers. Vision-Language Model (VLM) is used to parse scene information by image inputs to achieve a human-level understanding of terrain complexity. Based on the complexity classification, the system switches to the most suitable navigation mode, composing of perception, mapping and planning modules designed for different terrain types, to traverse the terrain ahead before reaching the next waypoint. By integrating the local navigation system with a map server and a global waypoint generation module, the rover is equipped to handle long-distance navigation tasks in complex scenarios. The navigation system is evaluated in various simulation environments. Compared to the single-mode conservative navigation method, our multi-mode system is able to bootstrap the time and energy efficiency in a long-distance traversal with varied type of obstacles, enhancing efficiency by 79.5%, while maintaining its avoidance capabilities against terrain hazards to guarantee rover safety. More system information is shown at https://chengsn1234.github.io/multi-mode-planetary-navigation/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。