arXiv:2507.18111cs.LG2025-07

用分位数强化学习优化开放基站切片,降低延迟并提升资源利用率。

Percentile-Based Deep Reinforcement Learning and Reward Based Personalization For Delay Aware RAN Slicing in O-RAN

  • 基于大数定律设计奖励函数,结合分位数实现延迟感知的强化学习。
  • 相比平均延迟优化模型,平均延迟降低38%,满足概率延迟约束。
  • 通过奖励驱动个性化权重共享,适合多运营商协同部署场景。

本文针对开放无线接入网(O-RAN)架构下的无线接入网(RAN)切片挑战,研究多个移动虚拟网络运营商(MVNOs)竞争物理资源块(PRBs)的场景,目标是在满足客户端概率延迟上限约束的前提下最小化PRB利用率。首先基于大数定律推导奖励函数,并针对实际实验环境进行实用化修改。提出分位数感知的延迟感知深度强化学习(PDA-DRL)方法,在多个基线中表现更优,实现38%的平均延迟下降。进一步研究多MVNO间模型权重共享问题,提出基于奖励的个性化方法:各智能体根据其他代理的表现优先级选择其模型权重。该方法优于传统的联邦平均聚合、依赖流量模式或权重距离相似性的策略。

原文摘要 · Abstract (English)

In this paper, we tackle the challenge of radio access network (RAN) slicing within an open RAN (O-RAN) architecture. Our focus centers on a network that includes multiple mobile virtual network operators (MVNOs) competing for physical resource blocks (PRBs) with the goal of meeting probabilistic delay upper bound constraints for their clients while minimizing PRB utilization. Initially, we derive a reward function based on the law of large numbers (LLN), then implement practical modifications to adapt it for real-world experimental scenarios. We then propose our solution, the Percentile-based Delay-Aware Deep Reinforcement Learning (PDA-DRL), which demonstrates its superiority over several baselines, including DRL models optimized for average delay constraints, by achieving a 38\% reduction in resultant average delay. Furthermore, we delve into the issue of model weight sharing among multiple MVNOs to develop a robust personalized model. We introduce a reward-based personalization method where each agent prioritizes other agents' model weights based on their performance. This technique surpasses traditional aggregation methods, such as federated averaging, and strategies reliant on traffic patterns and model weight distance similarities.

强化学习网络切片O-RAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。