让机器人通过看视频学会导航,跨平台通用。
LMD-PGN: Cross-Modal Knowledge Distillation from First-Person-View Images to Third-Person-View BEV Maps for Universal Point Goal Navigation
- 用第一视角图像和动作,蒸馏出通用的第三人称地图与子目标
- 在仿真环境中实现跨平台导航模型迁移,无需重新训练
- 适合多机器人协作、异构平台的智能系统部署
点目标导航(PGN)是一种无地图的导航方法,使机器人仅依赖视觉信息到达目标点。尽管深度强化学习已显著提升复杂环境下的导航能力,但现有方法仅针对单机器人系统,难以推广至多机器人、异构平台场景。本文提出一种知识迁移框架,使教师机器人将导航策略转移给学生机器人,包括未知或黑箱平台。引入新颖的知识蒸馏(KD)机制,将第一人称视角(FPV)表示(图像、转向/前进动作)转化为通用的第三人称视角(TPV)表示(局部地图、子目标)。状态重构为基于SLAM的局部地图,动作映射为预定义网格上的子目标。为提高训练效率,提出一种采样高效的KD方法,通过噪声鲁棒的局部地图描述符(LMD)对齐训练轨迹。实验在Habitat-Sim中验证了该框架的可行性,实施成本极低。研究展示了可扩展、跨平台的PGN解决方案潜力,拓展了具身AI在多机器人场景中的应用边界。
原文摘要 · Abstract (English)
Point goal navigation (PGN) is a mapless navigation approach that trains robots to visually navigate to goal points without relying on pre-built maps. Despite significant progress in handling complex environments using deep reinforcement learning, current PGN methods are designed for single-robot systems, limiting their generalizability to multi-robot scenarios with diverse platforms. This paper addresses this limitation by proposing a knowledge transfer framework for PGN, allowing a teacher robot to transfer its learned navigation model to student robots, including those with unknown or black-box platforms. We introduce a novel knowledge distillation (KD) framework that transfers first-person-view (FPV) representations (view images, turning/forward actions) to universally applicable third-person-view (TPV) representations (local maps, subgoals). The state is redefined as reconstructed local maps using SLAM, while actions are mapped to subgoals on a predefined grid. To enhance training efficiency, we propose a sampling-efficient KD approach that aligns training episodes via a noise-robust local map descriptor (LMD). Although validated on 2D wheeled robots, this method can be extended to 3D action spaces, such as drones. Experiments conducted in Habitat-Sim demonstrate the feasibility of the proposed framework, requiring minimal implementation effort. This study highlights the potential for scalable and cross-platform PGN solutions, expanding the applicability of embodied AI systems in multi-robot scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。