arXiv:2511.15002cs.AIcs.LG2025-11中稿 · be published in IE…被引 4

用智能调控的强化学习,让5G网络资源分配更高效稳定

Task Specific Sharpness Aware O-RAN Resource Management using Multi Agent Reinforcement Learning

  • 通过动态调整正则化强度,只对复杂环境中的代理进行优化
  • 资源分配效率提升22%,不同网络切片下服务质量均更优
  • 适合研究O-RAN智能调度、多智能体强化学习的工程师和学者

下一代网络采用开放无线接入网(O-RAN)架构实现动态资源管理,由无线接入网智能控制器(RIC)驱动。尽管深度强化学习(DRL)在优化网络资源方面展现潜力,但在动态环境中常面临鲁棒性与泛化能力不足的问题。本文提出一种新型资源管理方法,在分布式多智能体强化学习(MARL)框架中改进Soft Actor Critic(SAC)算法,引入尖锐度感知最小化(SAM),并设计自适应、选择性机制:正则化强度由时序差分(TD)误差方差决定,仅对环境复杂度高的智能体施加正则化。该策略有效降低冗余开销,提升训练稳定性与泛化性能,且不牺牲学习效率。同时引入动态ρ调度方案,优化各智能体的探索-利用平衡。实验表明,该方法显著优于传统DRL方法,在资源分配效率上最高提升22%,并在多种O-RAN切片场景下保障了更优的服务质量(QoS)满足度。

原文摘要 · Abstract (English)

Next-generation networks utilize the Open Radio Access Network (O-RAN) architecture to enable dynamic resource management, facilitated by the RAN Intelligent Controller (RIC). While deep reinforcement learning (DRL) models show promise in optimizing network resources, they often struggle with robustness and generalizability in dynamic environments. This paper introduces a novel resource management approach that enhances the Soft Actor Critic (SAC) algorithm with Sharpness-Aware Minimization (SAM) in a distributed Multi-Agent RL (MARL) framework. Our method introduces an adaptive and selective SAM mechanism, where regularization is explicitly driven by temporal-difference (TD)-error variance, ensuring that only agents facing high environmental complexity are regularized. This targeted strategy reduces unnecessary overhead, improves training stability, and enhances generalization without sacrificing learning efficiency. We further incorporate a dynamic $ρ$ scheduling scheme to refine the exploration-exploitation trade-off across agents. Experimental results show our method significantly outperforms conventional DRL approaches, yielding up to a $22\%$ improvement in resource allocation efficiency and ensuring superior QoS satisfaction across diverse O-RAN slices.

O-RAN强化学习资源管理多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。