用多智能体强化学习实现基站负载均衡的实时用户关联与切换。
Multi-Agent Q-Learning for Real-Time Load Balancing User Association and Handover in Mobile Networks
- 设计集中与分布式双策略,让用户和基站协同决策
- 在步行到高速驾驶下均实现低切换频率与高性能
- 适合动态蜂窝网络中需实时响应的场景
随着下一代蜂窝网络日益密集,如何在每一时刻将用户最优地关联至基站,同时避免基站过载,对保障网络稳定性和高性能至关重要。本文提出基于多智能体在线Q-learning(QL)算法,实现密集蜂窝网络中的实时负载均衡用户关联与切换。所有基站的负载约束使用户智能体的动作相互耦合,为此提出两种多智能体动作选择策略:一种是集中式,由中央负载平衡器(CLB)通过交换最差连接以最大化总学习奖励;另一种是分布式,每个用户设备(UE)基于本地信息参与与基站的分布式匹配游戏,以最大化局部奖励。将这两种策略集成至在线QL算法中,该算法可实时适应信道变化与用户移动性,采用考虑切换代价的奖励函数以减少切换频率。所提算法具备低复杂度与快速收敛特性,优于3GPP最大信干噪比(max-SINR)关联策略。两种策略在步行、跑步、骑车及郊区驾驶等多种用户速度场景下均表现良好,展现出强鲁棒性与实时适应能力。
原文摘要 · Abstract (English)
As next generation cellular networks become denser, associating users with the optimal base stations at each time while ensuring no base station is overloaded becomes critical for achieving stable and high network performance. We propose multi-agent online Q-learning (QL) algorithms for performing real-time load balancing user association and handover in dense cellular networks. The load balancing constraints at all base stations couple the actions of user agents, and we propose two multi-agent action selection policies, one centralized and one distributed, to satisfy load balancing at every learning step. In the centralized policy, the actions of UEs are determined by a central load balancer (CLB) running an algorithm based on swapping the worst connection to maximize the total learning reward. In the distributed policy, each UE takes an action based on its local information by participating in a distributed matching game with the BSs to maximize the local reward. We then integrate these action selection policies into an online QL algorithm that adapts in real-time to network dynamics including channel variations and user mobility, using a reward function that considers a handover cost to reduce handover frequency. The proposed multi-agent QL algorithm features low-complexity and fast convergence, outperforming 3GPP max-SINR association. Both policies adapt well to network dynamics at various UE speed profiles from walking, running, to biking and suburban driving, illustrating their robustness and real-time adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。