提出分层强化学习框架,让无人机群在通信受限下高效协同避障与追踪。
Decentralized Consensus Inference-based Hierarchical Reinforcement Learning for Multi-Constrained UAV Pursuit-Evasion Game
- 分两层设计:高层定位目标,底层控制避障与编队
- 通信受限下仍能实现高精度协同,任务完成率超90%
- 适合复杂动态环境中的多无人机协同任务
多旋翼无人机系统在多约束追逃游戏(MC-PEG)中备受关注,尤其在通信受限条件下,无人机集群需同时实现目标区域覆盖与协同避险。本文提出一种两级分层强化学习框架——共识推理型分层强化学习(CI-HRL),将目标定位交由高层策略,低层策略负责避障、导航与编队控制。高层采用新型多智能体通信模块ConsMAC,通过聚合邻近信息实现从局部状态到全局共识的感知。低层则结合交替训练的多智能体近端策略优化(AT-M)与策略蒸馏技术。软件在环(SITL)仿真验证表明,该方法在通信受限场景下显著提升集群协作避险与任务完成能力,任务成功率超过90%。
原文摘要 · Abstract (English)
Multiple quadrotor unmanned aerial vehicle (UAV) systems have garnered widespread research interest and fostered tremendous interesting applications, especially in multi-constrained pursuit-evasion games (MC-PEG). The Cooperative Evasion and Formation Coverage (CEFC) task, where the UAV swarm aims to maximize formation coverage across multiple target zones while collaboratively evading predators, belongs to one of the most challenging issues in MC-PEG, especially under communication-limited constraints. This multifaceted problem, which intertwines responses to obstacles, adversaries, target zones, and formation dynamics, brings up significant high-dimensional complications in locating a solution. In this paper, we propose a novel two-level framework (i.e., Consensus Inference-based Hierarchical Reinforcement Learning (CI-HRL)), which delegates target localization to a high-level policy, while adopting a low-level policy to manage obstacle avoidance, navigation, and formation. Specifically, in the high-level policy, we develop a novel multi-agent reinforcement learning module, Consensus-oriented Multi-Agent Communication (ConsMAC), to enable agents to perceive global information and establish consensus from local states by effectively aggregating neighbor messages. Meanwhile, we leverage an Alternative Training-based Multi-agent proximal policy optimization (AT-M) and policy distillation to accomplish the low-level control. The experimental results, including the high-fidelity software-in-the-loop (SITL) simulations, validate that CI-HRL provides a superior solution with enhanced swarm's collaborative evasion and task completion capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。