arXiv:2410.18495cs.RO2024-10

用强化学习让多架无人机在避障中保持队形

Multi-UAV Formation Control with Static and Dynamic Obstacle Avoidance via Reinforcement Learning

  • 分两阶段训练:先搜寻合理奖励函数,再用课程学习加速训练
  • 仿真与真实环境测试中,避障成功率和队形保持率均优于基线方法
  • 引入注意力编码器提升复杂障碍场景下的适应能力,适合无人机集群应用

本文解决多架无人机在定向飞行中同时维持队形并避开静态与动态障碍的难题。任务难点在于多目标平衡、探索空间大及仿真到现实的差距。为此,提出两阶段强化学习流程:第一阶段随机搜索兼顾定向飞行、避障、队形保持和零样本部署的奖励函数;第二阶段使用该奖励函数,在更复杂场景中结合课程学习加速策略训练。此外,引入基于注意力的观测编码器,提升队形维持能力与对不同障碍密度的适应性。仿真与真实环境实验表明,该方法在静态、动态及混合障碍场景下,碰撞规避率和队形保持率均优于基于规划和强化学习的基线方法。消融实验验证了课程学习策略与注意力编码器的有效性。动画演示见:https://sites.google.com/view/uav-formation-with-avoidance/

原文摘要 · Abstract (English)

This paper tackles the challenging task of maintaining formation among multiple unmanned aerial vehicles (UAVs) while avoiding both static and dynamic obstacles during directed flight. The complexity of the task arises from its multi-objective nature, the large exploration space, and the sim-to-real gap. To address these challenges, we propose a two-stage reinforcement learning (RL) pipeline. In the first stage, we randomly search for a reward function that balances key objectives: directed flight, obstacle avoidance, formation maintenance, and zero-shot policy deployment. The second stage applies this reward function to more complex scenarios and utilizes curriculum learning to accelerate policy training. Additionally, we incorporate an attention-based observation encoder to improve formation maintenance and adaptability to varying obstacle densities. Experimental results in both simulation and real-world environments demonstrate that our method outperforms both planning-based and RL-based baselines in terms of collision-free rates and formation maintenance across static, dynamic, and mixed obstacle scenarios. Ablation studies further confirm the effectiveness of our curriculum learning strategy and attention-based encoder. Animated demonstrations are available at: https://sites.google.com/view/ uav-formation-with-avoidance/.

无人机集群强化学习避障队形控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。