用图像生成模型让机器人看懂指令并导航160米远距离路线
PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation

- 用语言指令生成目标和可通行区域图,指导机器人规划路径
- 通过跨视角定位纠正里程计漂移,实现160米户外长距导航
- 适合需要强泛化能力的地面与无人机自主导航场景
鸟瞰图(BEV)被广泛证明对导航具有重要先验信息。尽管其提供全局视野,但如何充分挖掘该信息以及在执行中可靠使用仍是挑战。本文提出一种基于BEV图作为全局先验的导航系统,适用于地面及近地机器人平台。系统利用图像生成模型从自然语言中解析人类意图、识别目标位置,并生成可通行掩码。执行阶段引入跨视图定位,将机器人里程计与BEV地图对齐,有效缓解传统里程计的长期漂移问题。我们在多个基准上进行实验评估,并进一步在无人机平台上验证。仅使用常规局部运动规划器,无人机成功完成160米室外长距离导航任务。本工作展示了基础模型的世界理解能力如何迁移至具身导航,使机器人能受益于现有图像生成模型的强大泛化能力。
原文摘要 · Abstract (English)
Bird's-eye-view (BEV) images have been widely demonstrated to provide valuable prior information for navigation. Given the global information provided by such views, two key challenges remain: how to fully exploit this information and how to reliably use it during execution. In this paper, we propose a navigation system that uses BEV images as global priors and is designed for ground and near-ground robotic platforms. The system employs an image generation model to interpret human intent from natural language, identify the target destination, and generate traversability masks. During execution, we introduce cross-view localization to align the robot's odometry with the BEV map and mitigate long-term drift in conventional odometry. We conduct extensive benchmark experiments to evaluate the proposed method and further validate it on a UAV platform. Using only a conventional local motion planner, the UAV successfully completes a 160-meter outdoor long-range navigation task. This work demonstrates how the world-understanding capabilities of foundation models can be transferred to embodied navigation, enabling robots to benefit from the strong generalization ability of existing image generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。