用分层强化学习提升多无人机协同作战能力
A Hierarchical Reinforcement Learning Framework for Multi-UAV Combat Using Leader-Follower Strategy
- 分三层架构:宏观决策、角度选择、精确动作生成
- 采用领袖-跟随策略,显著提升多机协作效率
- 适合研究多智能体协同与无人机对抗的学者
多无人机空战是涉及多个自主无人机的复杂任务,是航空航天与人工智能交叉领域的前沿方向。现有方法大多将动作空间离散化为预定义动作,限制了无人机机动性与复杂策略实现;部分研究简化为1对1对抗,忽略了多机间的协作动态。为应对六自由度空间中的高维挑战并提升协作能力,本文提出一种基于领袖-跟随多智能体近端策略优化(LFMAPPO)的分层框架。该框架分为三级:顶层进行环境宏观评估并指导执行策略;中层决定期望动作的角度;底层生成高维动作空间中的精确指令。通过为不同角色分配状态价值函数,并利用领袖-跟随机制训练顶层策略,使跟随者估计领袖效用,促进智能体间有效协作。此外,引入与无人机姿态对齐的目标选择器,评估目标威胁等级。仿真实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Multi-UAV air combat is a complex task involving multiple autonomous UAVs, an evolving field in both aerospace and artificial intelligence. This paper aims to enhance adversarial performance through collaborative strategies. Previous approaches predominantly discretize the action space into predefined actions, limiting UAV maneuverability and complex strategy implementation. Others simplify the problem to 1v1 combat, neglecting the cooperative dynamics among multiple UAVs. To address the high-dimensional challenges inherent in six-degree-of-freedom space and improve cooperation, we propose a hierarchical framework utilizing the Leader-Follower Multi-Agent Proximal Policy Optimization (LFMAPPO) strategy. Specifically, the framework is structured into three levels. The top level conducts a macro-level assessment of the environment and guides execution policy. The middle level determines the angle of the desired action. The bottom level generates precise action commands for the high-dimensional action space. Moreover, we optimize the state-value functions by assigning distinct roles with the leader-follower strategy to train the top-level policy, followers estimate the leader's utility, promoting effective cooperation among agents. Additionally, the incorporation of a target selector, aligned with the UAVs' posture, assesses the threat level of targets. Finally, simulation experiments validate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。