用视觉导航无人机在未知环境自主避障,无需人工操控。
Vision-Guided Outdoor Flight and Obstacle Evasion via Reinforcement Learning

- 结合立体视觉与惯性定位,通过强化学习训练自主导航策略。
- 在仿真中分两阶段训练,实现零样本迁移至真实场景。
- 适合无人飞行器自主导航、野外搜救等复杂环境应用。
尽管四旋翼飞行器凭借全向机动能力具备出色的穿越能力,但在复杂环境中持续依赖飞行员操控限制了其在无GNSS和遥测信号场景的应用。为此,我们提出一种新型感知-运动策略,利用双目视觉深度和视觉惯性里程计(VIO)实现未知环境中自主导航至目标点并避障。该策略由预训练自编码器作为感知头,后接规划与控制的LSTM网络,输出可被商用无人机直接执行的速度指令。通过两阶段强化学习与特权学习训练:1)初始阶段使用全局运动规划器生成的最优轨迹作为监督指导;2)在课程化环境中进一步微调。为缩小仿真到现实的差距,采用领域随机化与奖励塑形,使策略对噪声和域偏移具有鲁棒性。户外实验表明,该方法成功实现零样本迁移至未训练过的障碍物环境及无人机平台。
原文摘要 · Abstract (English)
Although quadcopters boast impressive traversal capabilities enabled by their omnidirectional maneuverability, the need for continuous pilot control in complex environments impedes their application in GNSS and telemetry-denied scenarios. To this end, we propose a novel sensorimotor policy that uses stereo-vision depth and visual-inertial odometry (VIO) to autonomously navigate through obstacles in an unknown environment to reach a goal point. The policy is comprised of a pre-trained autoencoder as the perception head followed by a planning and control LSTM network which outputs velocity commands that can be followed by an off-the-shelf commercial drone. We leverage reinforcement and privileged learning paradigms to train the policy in simulation through a two-stage process: 1) initial training with optimal trajectories generated by a global motion planner acting as a supervisory backbone, 2) further fine-tuning in a curriculum environment. To bridge the sim-to-real gap, we employ domain randomization and reward shaping to create a policy that is both robust to noise and domain shift. In outdoor experiments, our approach achieves successful zero-shot transfer to both obstacle environments and a drone platform that were never encountered during training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。