用特权信息提升无人机绕大障碍物的强化学习导航能力
Quadrotor Navigation using Reinforcement Learning with Privileged Information
- 引入到达时间图作为特权信息,结合航向对齐损失指导飞行
- 在复杂场景中成功率达86%,比基线高出34%
- 实机测试20次共589米无碰撞,昼夜皆可稳定运行
本文提出一种基于强化学习的四旋翼无人机导航方法,利用高效可微分仿真、新颖损失函数及特权信息,在包含大障碍物、尖角和死胡同的拟真环境中实现导航。以往基于学习的方法在狭窄障碍物场景表现良好,但在目标被大型墙体或地形阻挡时失效。本方法引入到达时间(ToA)地图作为特权信息,并设计航向对齐损失,引导机器人绕过大障碍。策略在含大障碍的复杂场景中测试,成功率达到86%,优于基线34%。在真实室外杂乱环境中部署于定制四旋翼,完成20次飞行,总距离589米,最高速度达4米/秒,全程无碰撞。
原文摘要 · Abstract (English)
This paper presents a reinforcement learning-based quadrotor navigation method that leverages efficient differentiable simulation, novel loss functions, and privileged information to navigate around large obstacles. Prior learning-based methods perform well in scenes that exhibit narrow obstacles, but struggle when the goal location is blocked by large walls or terrain. In contrast, the proposed method utilizes time-of-arrival (ToA) maps as privileged information and a yaw alignment loss to guide the robot around large obstacles. The policy is evaluated in photo-realistic simulation environments containing large obstacles, sharp corners, and dead-ends. Our approach achieves an 86% success rate and outperforms baseline strategies by 34%. We deploy the policy onboard a custom quadrotor in outdoor cluttered environments both during the day and night. The policy is validated across 20 flights, covering 589 meters without collisions at speeds up to 4 m/s.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。