多智能体强化学习让无人机群自动避障抗干扰,吞吐量提升50%
Multi-Agent Deep Reinforcement Learning for Collaborative UAV Relay Networks under Jamming Atatcks
- 用集中训练分散执行框架,全局决策指导局部行动
- 吞吐量提升约50%,碰撞率接近零,抗干扰能力显著增强
- 无需预设抗干扰策略,智能体自发形成有效应对机制
将无人机群作为动态通信中继部署于下一代战术网络至关重要。然而,在对抗环境中需权衡多项目标:最大化系统吞吐量、避免碰撞并抵御敌方干扰。现有基于启发式的方法因问题的动态性和多目标特性难以取得理想解。本文将该问题建模为协作多智能体强化学习(MARL)任务,采用集中训练分散执行(CTDE)框架。中心化评价器利用全局状态信息指导去中心化执行者,后者仅依赖局部观测。仿真结果表明,所提框架显著优于启发式基线,总系统吞吐量提升约50%,同时碰撞率接近零。关键发现是,智能体在未显式编程的情况下自发发展出抗干扰策略,能智能定位以平衡干扰抑制与与地面用户通信链路的维持。
原文摘要 · Abstract (English)
The deployment of Unmanned Aerial Vehicle (UAV) swarms as dynamic communication relays is critical for next-generation tactical networks. However, operating in contested environments requires solving a complex trade-off, including maximizing system throughput while ensuring collision avoidance and resilience against adversarial jamming. Existing heuristic-based approaches often struggle to find effective solutions due to the dynamic and multi-objective nature of this problem. This paper formulates this challenge as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, solved using the Centralized Training with Decentralized Execution (CTDE) framework. Our approach employs a centralized critic that uses global state information to guide decentralized actors which operate using only local observations. Simulation results show that our proposed framework significantly outperforms heuristic baselines, increasing the total system throughput by approximately 50% while simultaneously achieving a near-zero collision rate. A key finding is that the agents develop an emergent anti-jamming strategy without explicit programming. They learn to intelligently position themselves to balance the trade-off between mitigating interference from jammers and maintaining effective communication links with ground users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。