arXiv:2505.08995cs.AIcs.LG2025-05被引 1

用分层强化学习模拟空战,提升复杂对抗策略生成能力

Enhancing Aerial Combat Tactics through Hierarchical Multi-Agent Reinforcement Learning

  • 分两层决策:底层控机,高层指挥全局任务
  • 通过渐进式训练,实现高阶作战策略有效生成
  • 适合军事仿真与智能战术系统研发人员参考

本文提出一种分层多智能体强化学习框架,用于分析涉及异构智能体的模拟空战场景。目标是在预设仿真中识别能达成任务成功的有效作战方案,从而以低成本、安全可试错的方式探索真实国防场景。在该背景下应用深度强化学习面临诸多挑战,包括复杂的飞行动力学、多智能体系统中状态与动作空间的指数级增长,以及实时单元控制与前瞻规划的融合难题。为此,决策过程被划分为两个抽象层级:底层策略负责单个单位的控制,高层指挥官策略则发出与整体任务目标对齐的宏观指令。该分层结构通过利用个体智能体策略的对称性,并将控制与命令任务分离,提升了训练效率。底层策略在逐步增加复杂度的课程中训练,以掌握个体作战控制;随后在预训练控制策略基础上训练高层指挥官策略,以实现任务目标。实证验证表明该框架具有显著优势。

原文摘要 · Abstract (English)

This work presents a Hierarchical Multi-Agent Reinforcement Learning framework for analyzing simulated air combat scenarios involving heterogeneous agents. The objective is to identify effective Courses of Action that lead to mission success within preset simulations, thereby enabling the exploration of real-world defense scenarios at low cost and in a safe-to-fail setting. Applying deep Reinforcement Learning in this context poses specific challenges, such as complex flight dynamics, the exponential size of the state and action spaces in multi-agent systems, and the capability to integrate real-time control of individual units with look-ahead planning. To address these challenges, the decision-making process is split into two levels of abstraction: low-level policies control individual units, while a high-level commander policy issues macro commands aligned with the overall mission targets. This hierarchical structure facilitates the training process by exploiting policy symmetries of individual agents and by separating control from command tasks. The low-level policies are trained for individual combat control in a curriculum of increasing complexity. The high-level commander is then trained on mission targets given pre-trained control policies. The empirical validation confirms the advantages of the proposed framework.

强化学习空战模拟分层决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。