用多智能体强化学习实现战场态势图的自动构建与抗干扰
Data-Driven Distributed Common Operational Picture from Heterogeneous Platforms using Multi-Agent Reinforcement Learning
- 通过强化学习让各平台自主编码感知信息并生成共享态势图
- 在星际争霸2中实现误差低于5%的精准态势图,通信中断下仍稳定运行
- 适合军事指挥系统、无人平台协同控制等高动态环境应用
配备先进传感器的无人平台集成有望提升军事行动中的态势感知能力,缓解“战争迷雾”。然而,海量异构数据给指挥控制(C2)系统带来巨大挑战。本文提出一种新型多智能体学习框架,实现智能体与人类之间的自主安全通信,实时构建可解释的共同作战图(COP)。每个智能体将感知与动作编码为紧凑向量,传输、接收并解码形成涵盖所有友敌方状态的全局态势图。采用深度强化学习(DRL)联合训练态势图模型与智能体动作策略。实验在星际争霸2仿真环境中验证,结果表明态势图误差低于5%,且在GPS失效、通信中断等恶劣条件下依然具备鲁棒性。本研究贡献包括:自主态势图生成方法、分布式预测增强的韧性,以及协同训练的态势图与多智能体强化学习策略。该工作推动了自适应、抗干扰的指挥控制发展,助力异构无人平台高效协同。
原文摘要 · Abstract (English)
The integration of unmanned platforms equipped with advanced sensors promises to enhance situational awareness and mitigate the "fog of war" in military operations. However, managing the vast influx of data from these platforms poses a significant challenge for Command and Control (C2) systems. This study presents a novel multi-agent learning framework to address this challenge. Our method enables autonomous and secure communication between agents and humans, which in turn enables real-time formation of an interpretable Common Operational Picture (COP). Each agent encodes its perceptions and actions into compact vectors, which are then transmitted, received and decoded to form a COP encompassing the current state of all agents (friendly and enemy) on the battlefield. Using Deep Reinforcement Learning (DRL), we jointly train COP models and agent's action selection policies. We demonstrate resilience to degraded conditions such as denied GPS and disrupted communications. Experimental validation is performed in the Starcraft-2 simulation environment to evaluate the precision of the COPs and robustness of policies. We report less than 5% error in COPs and policies resilient to various adversarial conditions. In summary, our contributions include a method for autonomous COP formation, increased resilience through distributed prediction, and joint training of COP models and multi-agent RL policies. This research advances adaptive and resilient C2, facilitating effective control of heterogeneous unmanned platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。