用多智能体强化学习优化校园毫米波基站部署,提升覆盖与公平性。
Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage

- 设计多智能体分布式策略,分区域协同优化基站位置。
- 在400用户场景下实现全覆盖,公平性指标达0.94。
- 相比单智能体,收敛更快,适合高密度复杂场景。
深度强化学习(DRL)因其实时自适应能力,在复杂优化问题中备受关注。本文研究真实非凸校园拓扑下的毫米波基站(BS)最优部署问题,该问题因最大最小公平性目标的非凸、非光滑特性而属于NP难问题。为此,将基站部署建模为马尔可夫决策过程(MDP),系统比较四种DRL方法:离散单智能体DQN、空间划分多智能体DQN、连续单智能体DDPG,以及地理分区多智能体DDPG框架。数值实验表明,多智能体DDPG在密集场景中显著优于单智能体方法;同时实现了全覆盖,并获得0.94的Jain公平性指数。此外,该方法在400用户密集场景下展现出高效的计算收敛性能。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) has recently gained a great attention due to its real-time adaptation and effectiveness in complex optimization problems. This paper investigates the optimal deployment of millimeter-wave (mmWave) base stations (BSs) in a realistic, non-convex campus topology. The optimization problem is NP-hard, due to the non-convex, non-smooth nature of the max-min fairness objective. To overcome these constraints, we formulate the BS placement as a Markov Decision Process (MDP) and systematically benchmark four DRL schemes: a discrete single-agent Deep Q-Network (DQN), a spatially partitioned Multi-Agent DQN, a continuous single-agent Deep Deterministic Policy Gradient (DDPG), and a geographically partitioned multi-agent DDPG framework. Numerical evaluations reveal that the multi-agent DDPG approach substantially outperforms single-agent in dense scenarios. Additionally full coverage is achieved, and a fairness Jain's index of 0.94 is obtained. Finally, the multi-agent demonstrates highly efficient computational convergence of dense scenarios with $400$ users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。