arXiv:2412.16167cs.NIcs.LG2024-12被引 15

用分层强化学习动态调整无人机集群,省电又稳。

Hierarchical Multi-Agent DRL Based Dynamic Cluster Reconfiguration for UAV Mobility Management

  • 分层多智能体强化学习,高层管集群、底层调功率。
  • 集群重配置频率降低,功率消耗减少,性能接近中心化方案。
  • 适合大规模无人机网络,扩展性好,部署灵活。

多连接性涉及分布式接入点(APs)间的动态集群形成及协同资源分配,对具有多连接性的用户提出高效移动性管理需求。本文提出一种面向无人机(UAVs)的新型移动性管理方案,通过无线干扰网络中的能量高效功率分配实现动态集群重构。目标包括满足严格的可靠性要求、最小化联合功率消耗以及降低集群重构频率。为此,我们设计了专用于动态聚类与功率分配的分层多智能体深度强化学习(H-MADRL)框架:边缘云通过低时延光回传链路连接一组AP,其上的高层智能体负责最优集群策略;低层智能体位于各AP中,负责功率分配策略。为提升学习效率,提出一种新的动作-观测过渡驱动学习算法,使低层智能体可将高层智能体的动作空间作为局部观测的一部分,从而共享部分集群策略信息,更高效地分配功率。仿真结果表明,所提分布式算法性能接近集中式算法;当AP数量翻倍时,集群与功率分配决策时间仅增加10%,而集中式方法增加90%。

原文摘要 · Abstract (English)

Multi-connectivity involves dynamic cluster formation among distributed access points (APs) and coordinated resource allocation from these APs, highlighting the need for efficient mobility management strategies for users with multi-connectivity. In this paper, we propose a novel mobility management scheme for unmanned aerial vehicles (UAVs) that uses dynamic cluster reconfiguration with energy-efficient power allocation in a wireless interference network. Our objective encompasses meeting stringent reliability demands, minimizing joint power consumption, and reducing the frequency of cluster reconfiguration. To achieve these objectives, we propose a hierarchical multi-agent deep reinforcement learning (H-MADRL) framework, specifically tailored for dynamic clustering and power allocation. The edge cloud connected with a set of APs through low latency optical back-haul links hosts the high-level agent responsible for the optimal clustering policy, while low-level agents reside in the APs and are responsible for the power allocation policy. To further improve the learning efficiency, we propose a novel action-observation transition-driven learning algorithm that allows the low-level agents to use the action space from the high-level agent as part of the local observation space. This allows the lower-level agents to share partial information about the clustering policy and allocate the power more efficiently. The simulation results demonstrate that our proposed distributed algorithm achieves comparable performance to the centralized algorithm. Additionally, it offers better scalability, as the decision time for clustering and power allocation increases by only 10% when doubling the number of APs, compared to a 90% increase observed with the centralized approach.

无人机强化学习集群管理分布式系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。