离线多智能体强化学习优化基站调度,提升用户速率与网络效率。
An Offline Multi-Agent Reinforcement Learning Framework for Radio Resource Management
- 基于离线多智能体强化学习,不依赖实时交互训练调度策略。
- 相比传统方法,综合速率提升超15%,尾部用户性能显著改善。
- 采用中心化训练、去中心化执行框架,兼顾效率与可扩展性。
离线多智能体强化学习(MARL)解决了在线MARL存在的安全风险、数据采集成本高、训练周期长和信令开销大等问题。本文提出一种面向无线资源管理(RRM)的离线MARL算法,旨在优化多个接入点(APs)的调度策略,联合最大化用户设备(UEs)的总速率与尾部速率。评估了三种训练范式:集中式、独立式以及中心化训练去中心化执行(CTDE)。仿真结果表明,所提离线MARL框架优于传统基线方法,在总速率与尾部速率加权组合上提升超过15%。其中,CTDE框架在降低集中式方法计算复杂度的同时,克服了独立训练的低效问题。这些结果证明了离线MARL在动态无线网络中实现可扩展、鲁棒且高效的资源管理解决方案的巨大潜力。
原文摘要 · Abstract (English)
Offline multi-agent reinforcement learning (MARL) addresses key limitations of online MARL, such as safety concerns, expensive data collection, extended training intervals, and high signaling overhead caused by online interactions with the environment. In this work, we propose an offline MARL algorithm for radio resource management (RRM), focusing on optimizing scheduling policies for multiple access points (APs) to jointly maximize the sum and tail rates of user equipment (UEs). We evaluate three training paradigms: centralized, independent, and centralized training with decentralized execution (CTDE). Our simulation results demonstrate that the proposed offline MARL framework outperforms conventional baseline approaches, achieving over a 15\% improvement in a weighted combination of sum and tail rates. Additionally, the CTDE framework strikes an effective balance, reducing the computational complexity of centralized methods while addressing the inefficiencies of independent training. These results underscore the potential of offline MARL to deliver scalable, robust, and efficient solutions for resource management in dynamic wireless networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。