用平均奖励强化学习提升无线资源管理效率
Average Reward Reinforcement Learning for Wireless Radio Resource Management
- 提出平均奖励框架下的新算法ARO SAC,适配无线网络长期优化目标
- 仿真显示系统性能比传统方法提升15%
- 适合研究无线资源管理与强化学习结合的学者和工程师
本文针对强化学习(RL)在无线通信资源管理(RRM)应用中一个关键却被忽视的问题展开研究:折扣奖励RL框架与无线网络优化的非折扣目标之间存在不匹配。据我们所知,这是首次系统性探讨该差异的工作,从问题建模出发,通过仿真量化了两者之间的差距。为弥合这一鸿沟,我们引入平均奖励强化学习,并提出一种名为平均奖励离策略软演员-评论家(ARO SAC)的新方法,它是经典软演员-评论家算法在平均奖励框架下的适配版本。仿真结果表明,该方法相较传统折扣奖励RL方法实现系统性能15%的提升,凸显了平均奖励强化学习在提高无线网络优化效率与有效性方面的潜力。
原文摘要 · Abstract (English)
In this paper, we address a crucial but often overlooked issue in applying reinforcement learning (RL) to radio resource management (RRM) in wireless communications: the mismatch between the discounted reward RL formulation and the undiscounted goal of wireless network optimization. To the best of our knowledge, we are the first to systematically investigate this discrepancy, starting with a discussion of the problem formulation followed by simulations that quantify the extent of the gap. To bridge this gap, we introduce the use of average reward RL, a method that aligns more closely with the long-term objectives of RRM. We propose a new method called the Average Reward Off policy Soft Actor Critic (ARO SAC) is an adaptation of the well known Soft Actor Critic algorithm in the average reward framework. This new method achieves significant performance improvement our simulation results demonstrate a 15% gain in the system performance over the traditional discounted reward RL approach, underscoring the potential of average reward RL in enhancing the efficiency and effectiveness of wireless network optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。