提出分层元强化学习框架,提升O-RAN网络资源管理效率与适应性。
Meta Hierarchical Reinforcement Learning for Scalable Resource Management in O-RAN
- 分层结构:高层统筹切片资源,底层实现切片内调度。
- 元更新按时序误差方差加权,提升复杂场景稳定性。
- 适合大规模动态无线网络,尤其对eMBB/URLLC/mMTC多业务场景有效。
现代应用对无线网络的实时自适应与高效资源管理能力提出更高要求。开放无线接入网(O-RAN)架构通过其射频智能控制器(RIC)模块,成为动态资源管理与网络切片的关键解决方案。尽管人工智能方法展现潜力,但多数方法在不可预测、高度动态条件下性能下降。本文提出一种受模型无关元学习(MAML)启发的自适应元分层强化学习(Meta-HRL)框架,联合优化O-RAN中的资源分配与网络切片。该框架融合分层控制与元学习,实现全局与局部适应:高层控制器跨切片分配资源,底层智能体执行切片内调度。自适应元更新机制根据时序差分误差方差对任务加权,增强稳定性并优先处理复杂网络场景。理论分析证明了两级学习过程的次线性收敛性与后悔率保证。仿真结果表明,相比基线强化学习与元强化学习方法,网络管理效率提升19.8%,适应速度更快,且在eMBB、URLLC和mMTC切片中均实现更高的服务质量满意度。消融实验与可扩展性研究进一步验证方法鲁棒性,在网络规模扩大时达到最高40%的加速适应,同时保持公平性、延迟与吞吐量的一致表现。
原文摘要 · Abstract (English)
The increasing complexity of modern applications demands wireless networks capable of real time adaptability and efficient resource management. The Open Radio Access Network (O-RAN) architecture, with its RAN Intelligent Controller (RIC) modules, has emerged as a pivotal solution for dynamic resource management and network slicing. While artificial intelligence (AI) driven methods have shown promise, most approaches struggle to maintain performance under unpredictable and highly dynamic conditions. This paper proposes an adaptive Meta Hierarchical Reinforcement Learning (Meta-HRL) framework, inspired by Model Agnostic Meta Learning (MAML), to jointly optimize resource allocation and network slicing in O-RAN. The framework integrates hierarchical control with meta learning to enable both global and local adaptation: the high-level controller allocates resources across slices, while low level agents perform intra slice scheduling. The adaptive meta-update mechanism weights tasks by temporal difference error variance, improving stability and prioritizing complex network scenarios. Theoretical analysis establishes sublinear convergence and regret guarantees for the two-level learning process. Simulation results demonstrate a 19.8% improvement in network management efficiency compared with baseline RL and meta-RL approaches, along with faster adaptation and higher QoS satisfaction across eMBB, URLLC, and mMTC slices. Additional ablation and scalability studies confirm the method's robustness, achieving up to 40% faster adaptation and consistent fairness, latency, and throughput performance as network scale increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。