用多视角视觉语言动作模型,让机器人学会复杂导航。
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- 基于360度视觉输入,用教师-学生框架融合多个专家经验
- 在合成环境中实现强泛化,性能超过训练它的强化学习老师
- 适合需要多场景鲁棒导航的机器人研究与应用
视觉导航政策被视为有前景的方向,因其模仿人类通过自身视角视觉观测进行导航。然而,视觉观测的光学信息难以像激光雷达点云或深度图那样显式建模,因此需要智能模型和大规模数据。为此,我们提出利用视觉-语言-动作(VLA)模型,以师生方式从合成专家数据中学习多样化的导航能力。具体地,我们构建了基于预训练大语言模型和视觉基础模型的多视角VLA模型MM-Nav(具备360度观测)。针对大规模导航数据,我们在三个定制化挑战环境中,通过三个分别具备路径到达、狭窄空间穿行和障碍规避能力的强化学习(RL)专家生成专家数据。我们在线收集数据并迭代训练该VLA模型,训练比例根据各能力表现动态平衡。在合成环境中的大量实验表明,该模型具备强大的泛化能力;更关键的是,其学生模型性能优于原始的强化学习教师模型,体现了多能力融合的协同效应。真实世界实验进一步验证了方法的有效性。
原文摘要 · Abstract (English)
Visual navigation policy is widely regarded as a promising direction, as it mimics humans by using egocentric visual observations for navigation. However, optical information of visual observations is difficult to be explicitly modeled like LiDAR point clouds or depth maps, which subsequently requires intelligent models and large-scale data. To this end, we propose to leverage the intelligence of the Vision-Language-Action (VLA) model to learn diverse navigation capabilities from synthetic expert data in a teacher-student manner. Specifically, we implement the VLA model, MM-Nav, as a multi-view VLA (with 360 observations) based on pretrained large language models and visual foundation models. For large-scale navigation data, we collect expert data from three reinforcement learning (RL) experts trained with privileged depth information in three challenging tailor-made environments for different navigation capabilities: reaching, squeezing, and avoiding. We iteratively train our VLA model using data collected online from RL experts, where the training ratio is dynamically balanced based on performance on individual capabilities. Through extensive experiments in synthetic environments, we demonstrate that our model achieves strong generalization capability. Moreover, we find that our student VLA model outperforms the RL teachers, demonstrating the synergistic effect of integrating multiple capabilities. Extensive real-world experiments further confirm the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。