arXiv:2503.12753cs.NIcs.AI2025-03ICML被引 11

用安全强化学习实现符合SLA的O-RAN切片,保障沉浸式应用低延迟。

SafeSlice: Enabling SLA-Compliant O-RAN Slicing via Safe Deep Reinforcement Learning

  • 设计风险敏感奖励函数,满足切片累积延迟约束。
  • 引入监督学习安全层,实时修正资源分配动作,避免瞬时违规。
  • 在真实VR游戏流量下验证,显著降低延迟与资源消耗。

基于深度强化学习(DRL)的切片策略在模拟环境中表现优异,但在开放无线接入网(O-RAN)等物理系统中因仿真-现实差距面临挑战,难以保证服务等级协议(SLA)的合规性,尤其是沉浸式应用严格的延迟要求。部署的DRL切片代理可能在未见场景中做出损害性能的资源分配决策。本文提出SafeSlice,同时应对O-RAN切片的累积(轨迹级)和瞬时(状态级)延迟约束。通过设计基于S型函数的风险敏感奖励函数,反映切片的延迟需求;并构建监督学习成本模型作为安全层,将切片代理的资源分配动作投影至最近的安全动作,满足瞬时约束。实验涵盖多种服务,包括真实虚拟现实(VR)游戏流量,在极端及动态部署条件下评估性能。结果表明,SafeSlice相较基线可降低83.23%的平均累积延迟、93.24%的瞬时延迟违规,以及22.13%的资源消耗。结果还显示其对延迟阈值配置变化具有鲁棒性,契合O-RAN范式下运营商灵活配置的需求。

原文摘要 · Abstract (English)

Deep reinforcement learning (DRL)-based slicing policies have shown significant success in simulated environments but face challenges in physical systems such as open radio access networks (O-RANs) due to simulation-to-reality gaps. These policies often lack safety guarantees to ensure compliance with service level agreements (SLAs), such as the strict latency requirements of immersive applications. As a result, a deployed DRL slicing agent may make resource allocation (RA) decisions that degrade system performance, particularly in previously unseen scenarios. Real-world immersive applications require maintaining SLA constraints throughout deployment to prevent risky DRL exploration. In this paper, we propose SafeSlice to address both the cumulative (trajectory-wise) and instantaneous (state-wise) latency constraints of O-RAN slices. We incorporate the cumulative constraints by designing a sigmoid-based risk-sensitive reward function that reflects the slices' latency requirements. Moreover, we build a supervised learning cost model as part of a safety layer that projects the slicing agent's RA actions to the nearest safe actions, fulfilling instantaneous constraints. We conduct an exhaustive experiment that supports multiple services, including real virtual reality (VR) gaming traffic, to investigate the performance of SafeSlice under extreme and changing deployment conditions. SafeSlice achieves reductions of up to 83.23% in average cumulative latency, 93.24% in instantaneous latency violations, and 22.13% in resource consumption compared to the baselines. The results also indicate SafeSlice's robustness to changing the threshold configurations of latency constraints, a vital deployment scenario that will be realized by the O-RAN paradigm to empower mobile network operators (MNOs).

O-RAN强化学习网络切片低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。