arXiv:2412.01839cs.NIcs.LG2024-12被引 4

对比PPO与ACER在O-RAN资源分配中的表现,验证其优劣。

Dynamics of Resource Allocation in O-RANs: An In-depth Exploration of On-Policy and Off-Policy Deep Reinforcement Learning for Real-Time Applications

  • 用PPO(在线)和ACER(离线)模型解决O-RAN实时资源分配问题。
  • PPO平衡能效与延迟,ACER收敛更快,均优于贪心算法。
  • 适合5G/O-RAN系统优化与机器学习落地研究者参考。

深度强化学习(DRL)是解决移动网络复杂问题的强大工具。本文研究了两种DRL模型——基于策略的PPO与非基于策略的ACER——在开放无线接入网(O-RAN)资源分配中的应用。针对对服务质量(QoS)有严格要求的场景,研究在相同实验设置下对比了两类模型在时延敏感与容忍用户场景下的性能。本研究基于Nessrine Hammami与Kim Khoa Nguyen的原始工作进行复现,旨在验证并证明其结论。结果表明,两类DRL模型均优于贪心算法;其中PPO在能耗与用户延迟间取得更好平衡,而ACER表现出更快速的收敛速度。该分析支持了原始研究的可复现性与普适性,为优化O-RAN资源分配策略提供了重要参考。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) is a powerful tool used for addressing complex challenges in mobile networks. This paper investigates the application of two DRL models, on-policy and off-policy, in the field of resource allocation for Open Radio Access Networks (O-RAN). The on-policy model is the Proximal Policy Optimization (PPO), and the off-policy model is the Sample Efficient Actor-Critic with Experience Replay (ACER), which focuses on resolving the challenges of resource allocation associated with a Quality of Service (QoS) application that has strict requirements. Motivated by the original work of Nessrine Hammami and Kim Khoa Nguyen, this study is a replication to validate and prove the findings. Both PPO and ACER are used within the same experimental setup to assess their performance in a scenario of latency-sensitive and latency-tolerant users and compare them. The aim is to verify the efficacy of on-policy and off-policy DRL models in the context of O-RAN resource allocation. Results from this replication contribute to the ongoing scientific research and offer insights into the reproducibility and generalizability of the original research. This analysis reaffirms that both on-policy and off-policy DRL models have better performance than greedy algorithms in O-RAN settings. In addition, it confirms the original observations that the on-policy model (PPO) gives a favorable balance between energy consumption and user latency, while the off-policy model (ACER) shows a faster convergence. These findings give good insights to optimize resource allocation strategies in O-RANs. Index Terms: 5G, O-RAN, resource allocation, ML, DRL, PPO, ACER.

O-RAN强化学习资源分配5G

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。