提出离线分布强化学习,无需环境交互就能优化无线资源管理。
Offline and Distributional Reinforcement Learning for Radio Resource Management
- 用静态数据集离线训练,不依赖实时环境交互
- 通过回报分布建模不确定性,性能优于传统方法10%
- 适合无法在线试错的现实无线网络场景
强化学习(RL)在未来的智能无线网络中展现出巨大潜力。尽管在线RL已被用于无线资源管理(RRM),取代传统方案,但其依赖与环境的实时交互,在无法实现在线交互的实际问题中作用受限。此外,传统RL难以应对真实随机环境中存在的不确定性与风险。为此,本文提出一种离线且分布式的强化学习方案用于RRM问题,可在无环境交互的情况下使用静态数据集进行训练,并通过回报分布来考虑不确定性来源。仿真结果表明,该方案优于传统资源管理模型,且是唯一一个性能超过在线RL的方法,相比在线RL提升了10%。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has proved to have a promising role in future intelligent wireless networks. Online RL has been adopted for radio resource management (RRM), taking over traditional schemes. However, due to its reliance on online interaction with the environment, its role becomes limited in practical, real-world problems where online interaction is not feasible. In addition, traditional RL stands short in front of the uncertainties and risks in real-world stochastic environments. In this manner, we propose an offline and distributional RL scheme for the RRM problem, enabling offline training using a static dataset without any interaction with the environment and considering the sources of uncertainties using the distributions of the return. Simulation results demonstrate that the proposed scheme outperforms conventional resource management models. In addition, it is the only scheme that surpasses online RL with a 10 % gain over online RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。