用强化学习让多架无人机自主协作追逃,提升效率与协同能力。
Autonomous Decision Making for UAV Cooperative Pursuit-Evasion Game with Reinforcement Learning
- 基于异步双深度Q网络与优先经验回放,高效训练高维状态下的决策策略。
- 在不同无人机数量和任务场景下,实现角色与目标的动态分配,显著提升合作效率。
- 适用于复杂空战环境中的多机协同智能决策,适合无人系统研究者参考。
无人机智能决策应用日益广泛,1对1追逃博弈发展后,多无人机协同博弈成为新挑战。本文提出一种基于深度强化学习的多角色无人机协同追逃决策模型,以应对复杂环境下无人机自主决策难题。为提升强化学习在高维状态-动作空间中的训练效率,提出多环境异步双深度Q网络结合优先经验回放算法,有效训练无人机博弈策略。此外,为增强协作能力、提高任务完成效率并降低无人机代价,本文聚焦多无人机环境中角色与目标的分配。通过在不同场景下为无人机分配多样化任务与角色,获得适应不同数量无人机的协同博弈决策模型。仿真结果表明,所提方法能有效实现无人机在追逃场景中的自主决策,并展现出显著的协作能力。
原文摘要 · Abstract (English)
The application of intelligent decision-making in unmanned aerial vehicle (UAV) is increasing, and with the development of UAV 1v1 pursuit-evasion game, multi-UAV cooperative game has emerged as a new challenge. This paper proposes a deep reinforcement learning-based model for decision-making in multi-role UAV cooperative pursuit-evasion game, to address the challenge of enabling UAV to autonomously make decisions in complex game environments. In order to enhance the training efficiency of the reinforcement learning algorithm in UAV pursuit-evasion game environment that has high-dimensional state-action space, this paper proposes multi-environment asynchronous double deep Q-network with priority experience replay algorithm to effectively train the UAV's game policy. Furthermore, aiming to improve cooperation ability and task completion efficiency, as well as minimize the cost of UAVs in the pursuit-evasion game, this paper focuses on the allocation of roles and targets within multi-UAV environment. The cooperative game decision model with varying numbers of UAVs are obtained by assigning diverse tasks and roles to the UAVs in different scenarios. The simulation results demonstrate that the proposed method enables autonomous decision-making of the UAVs in pursuit-evasion game scenarios and exhibits significant capabilities in cooperation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。