用进化算法提升无人机协同空战的决策能力
Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat

- 融合遗传算法与强化学习,增强探索效率和多样性
- 在复杂对抗环境中实现更高胜率和更快收敛
- 适合研究自主无人机协同作战的学者与工程师
随着现代空战向超视距多机协同方向发展,无人作战飞机(UCAV)的自主决策面临高维状态空间、离散动作指令及强对抗动态环境的挑战。为克服现有多智能体强化学习方法在探索效率低、样本利用率差、策略泛化性弱等方面的局限,本文提出对抗式课程与进化增强的多智能体近端策略优化框架(ACE-MAPPO),通过引入遗传软更新机制提升种群多样性,避免陷入局部最优;采用进化增强的优先轨迹回放策略,提高稀疏高价值样本的利用效率;设计对抗式进化课程学习机制,实现难度递增的自适应训练。大量实验表明,该方法在训练稳定性、收敛速度和胜率方面均优于MAPPO及其他基线算法,验证了其在多机协同空战场景中的有效性。
原文摘要 · Abstract (English)
As modern air combat evolves toward beyond-visual-range (BVR) multi-aircraft cooperative engagements, autonomous decision-making for unmanned combat aerial vehicles (UCAVs) faces significant challenges due to high-dimensional state spaces, discrete action commands, and strongly adversarial dynamic environments. To overcome the limitations of existing multi-agent reinforcement learning (MARL) methods in such settings, namely insufficient exploration efficiency, low sample utilization, and poor policy generalization, we propose Adversarial Curriculum and Evolutionary-enhanced Multi-agent Proximal Policy Optimization (ACE-MAPPO), a hybrid learning framework that integrates evolutionary algorithms with MAPPO. Specifically, a genetic soft update mechanism is introduced to enhance population diversity and mitigate convergence to local optima. An evolutionary-augmented prioritized trajectory replay strategy is further employed to improve the utilization of sparse high-value samples. In addition, an adversarial evolutionary curriculum learning mechanism is designed to enable adaptive training with progressively increasing difficulty. Extensive experimental results demonstrate that the proposed method outperforms MAPPO and other baseline algorithms in terms of training stability, convergence speed, and win rate, validating its effectiveness in multi-aircraft cooperative air combat scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。