arXiv:2603.24634cs.NIcs.AI2026-03

用多智能体强化学习优化蜂窝网络切换参数,提升吞吐量并适应动态变化。

Dual-Graph Multi-Agent Reinforcement Learning for Handover Optimization

  • 将切换参数建模为双图上的分布式强化学习问题,每个智能体控制一对邻区偏移量
  • 在真实网络仿真中,相比传统规则和集中式强化学习,吞吐量显著提升
  • 适用于大规模移动网络,对拓扑和流量变化有强鲁棒性,适合运营商部署

蜂窝网络中的切换(HO)控制依赖于一系列通过规则启发式配置的参数。关键参数之一是小区个体偏移量(CIO),为每对邻近小区定义,用于调整切换触发决策。在网络规模下,调节CIO成为一个紧密耦合的问题:微小变动可能重定向多个邻区间的移动流量,而静态规则在非平稳流量与移动性条件下常失效。本文利用CIO的成对结构,将切换优化建模为网络双图上的分布式部分可观测马尔可夫决策过程(Dec-POMDP)。在此框架中,每个智能体控制一对邻区的CIO,观测其局部双图邻域内聚合的关键性能指标(KPI),实现可扩展的分布式决策并保持图局部性。基于此,提出TD3-D-MA,一种基于TD3的离散多智能体强化学习算法,采用共享参数图神经网络(GNN)作为策略网络,运行在双图上,并使用区域化双评判器进行训练,增强密集部署下的信用分配能力。在搭载真实运营商参数、覆盖异构流量场景与网络拓扑的ns-3系统级仿真环境中评估,结果表明TD3-D-MA在吞吐量上优于标准切换启发式方法及集中式强化学习基线,且在拓扑与流量变化下具备良好泛化能力。

原文摘要 · Abstract (English)

HandOver (HO) control in cellular networks is governed by a set of HO control parameters that are traditionally configured through rule-based heuristics. A key parameter for HO optimization is the Cell Individual Offset (CIO), defined for each pair of neighboring cells and used to bias HO triggering decisions. At network scale, tuning CIOs becomes a tightly coupled problem: small changes can redirect mobility flows across multiple neighbors, and static rules often degrade under non-stationary traffic and mobility. We exploit the pairwise structure of CIOs by formulating HO optimization as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) on the network's dual graph. In this representation, each agent controls a neighbor-pair CIO and observes Key Performance Indicators (KPIs) aggregated over its local dual-graph neighborhood, enabling scalable decentralized decisions while preserving graph locality. Building on this formulation, we propose TD3-D-MA, a discrete Multi-Agent Reinforcement Learning (MARL) variant of the TD3 algorithm with a shared-parameter Graph Neural Network (GNN) actor operating on the dual graph and region-wise double critics for training, improving credit assignment in dense deployments. We evaluate TD3-D-MA in an ns-3 system-level simulator configured with real-world network operator parameters across heterogeneous traffic regimes and network topologies. Results show that TD3-D-MA improves network throughput over standard HO heuristics and centralized RL baselines, and generalizes robustly under topology and traffic shifts.

强化学习网络优化多智能体蜂窝网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。