用特权信息提升无人机在复杂环境下的自主导航能力
Vision-Based Deep Reinforcement Learning of UAV Autonomous Navigation Using Privileged Information
- 训练时引入特权信息增强感知,解决观测不全问题
- 多智能体探索加速经验收集,提升收敛速度
- 在多种场景下表现更优,适合高动态飞行任务
无人机在复杂未知环境中的高效自主导航与避障能力,对农业灌溉、灾害救援和物流配送等应用至关重要。本文提出一种端到端的分布式特权强化学习(DPRL)导航算法,应对部分可观测条件下高速自主飞行的挑战。方法结合深度强化学习与特权学习,利用非对称的演员-评论家架构,在训练阶段向智能体提供特权信息,增强其感知能力。同时设计跨多样环境的多智能体探索策略,加速经验积累,促进模型快速收敛。我们在多种场景下开展大量仿真测试,将DPRL与当前最优导航算法对比,结果一致显示该算法在飞行效率、鲁棒性和整体成功率方面均表现更优。
原文摘要 · Abstract (English)
The capability of UAVs for efficient autonomous navigation and obstacle avoidance in complex and unknown environments is critical for applications in agricultural irrigation, disaster relief and logistics. In this paper, we propose the DPRL (Distributed Privileged Reinforcement Learning) navigation algorithm, an end-to-end policy designed to address the challenge of high-speed autonomous UAV navigation under partially observable environmental conditions. Our approach combines deep reinforcement learning with privileged learning to overcome the impact of observation data corruption caused by partial observability. We leverage an asymmetric Actor-Critic architecture to provide the agent with privileged information during training, which enhances the model's perceptual capabilities. Additionally, we present a multi-agent exploration strategy across diverse environments to accelerate experience collection, which in turn expedites model convergence. We conducted extensive simulations across various scenarios, benchmarking our DPRL algorithm against the state-of-the-art navigation algorithms. The results consistently demonstrate the superior performance of our algorithm in terms of flight efficiency, robustness and overall success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。