用图像处理知识设计奖励机制,让无人机群高效避障。
A Learning Framework For Cooperative Collision Avoidance of UAV Swarms Leveraging Domain Knowledge
- 基于图像处理的奖励机制,将障碍物建模为场中的峰值,自然避免碰撞。
- 支持大规模无人机群训练,无需复杂通信或信用分配机制。
- 可适应复杂环境,即使轮廓不可行也能通过训练自适应应对。
本文提出一种多智能体强化学习(MARL)框架,用于无人机群的协同避障,其奖励机制基于图像处理领域的先验知识,通过在二维场中近似轮廓实现。障碍物被建模为场中的极大值点,由于等高线不会穿过峰顶或相交,碰撞被天然规避。同时,该方法生成的轨迹平滑且能耗低。本框架支持大规模无人机群训练,因智能体间交互最小化,无需复杂信用分配或观测共享机制。此外,通过密集训练,无人机群可适应轮廓不可行或不存在的复杂环境。大量实验表明,该框架在性能上优于当前主流MARL算法。
原文摘要 · Abstract (English)
This paper presents a multi-agent reinforcement learning (MARL) framework for cooperative collision avoidance of UAV swarms leveraging domain knowledge-driven reward. The reward is derived from knowledge in the domain of image processing, approximating contours on a two-dimensional field. By modeling obstacles as maxima on the field, collisions are inherently avoided as contours never go through peaks or intersect. Additionally, counters are smooth and energy-efficient. Our framework enables training with large swarm sizes as the agent interaction is minimized and the need for complex credit assignment schemes or observation sharing mechanisms in state-of-the-art MARL approaches are eliminated. Moreover, UAVs obtain the ability to adapt to complex environments where contours may be non-viable or non-existent through intensive training. Extensive experiments are conducted to evaluate the performances of our framework against state-of-the-art MARL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。