用强化学习让无人机群边飞行边通信,省电又公平。
Integrated Communication and Control for Energy-Efficient UAV Swarms: A Multi-Agent Reinforcement Learning Approach
- 将无人机飞行与通信联合优化,用强化学习实时协调
- 通信公平性达0.99,能耗降低25%以上
- 适合应急救援、野外覆盖等低能量场景
无人机群辅助通信网络在基础设施薄弱区域的覆盖补盲中日益重要,尤其适用于应急救援、军事行动和偏远地区覆盖等临时场景。然而复杂地理环境导致无线信道动态变化,频繁中断空地链路,严重影响通信可靠性与服务质量。为提升复杂环境下的无人机群通信质量,本文提出一种通信与控制一体化协同设计机制。针对无人机群严苛的能量约束,该机制在保障移动地面用户通信速率公平性的同时,优化能源效率。将联合资源分配与三维轨迹控制问题建模为马尔可夫决策过程(MDP),并设计多智能体强化学习(MARL)框架实现无人机群实时协同。提出新型多智能体混合近端策略优化带动作掩码(MAHPPO-AM)算法,有效处理高维混合动作空间,并通过动作掩码强制执行硬约束。实验表明,该方法在保证公平性指数0.99的同时,相比基线方法能耗降低最高达25%。
原文摘要 · Abstract (English)
The deployment of unmanned aerial vehicle (UAV) swarm-assisted communication networks has become an increasingly vital approach for remediating coverage limitations in infrastructure-deficient environments, with especially pressing applications in temporary scenarios, such as emergency rescue, military and security operations, and remote area coverage. However, complex geographic environments lead to unpredictable and highly dynamic wireless channel conditions, resulting in frequent interruptions of air-to-ground (A2G) links that severely constrain the reliability and quality of service in UAV swarm-assisted mobile communications. To improve the quality of UAV swarm-assisted communications in complex geographic environments, we propose an integrated communication and control co-design mechanism. Given the stringent energy constraints inherent in UAV swarms, our proposed mechanism is designed to optimize energy efficiency while maintaining an equilibrium between equitable communication rates for mobile ground users (GUs) and UAV energy expenditure. We formulate the joint resource allocation and 3D trajectory control problem as a Markov decision process (MDP), and develop a multi-agent reinforcement learning (MARL) framework to enable real-time coordinated actions across the UAV swarm. To optimize the action policy of UAV swarms, we propose a novel multi-agent hybrid proximal policy optimization with action masking (MAHPPO-AM) algorithm, specifically designed to handle complex hybrid action spaces. The algorithm incorporates action masking to enforce hard constraints in high-dimensional action spaces. Experimental results demonstrate that our approach achieves a fairness index of 0.99 while reducing energy consumption by up to 25% compared to baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。