arXiv:2602.00035cs.NIcs.DC2026-02

异步多智能体强化学习实现5G路由高效调度

Asynchronous MultiAgent Reinforcement Learning for 5G Routing under Side Constraints

  • 每个服务独立运行PPO智能体,异步更新资源状态
  • 保持服务可用率与端到端延迟,训练时间减少
  • 适合高并发、动态变化的5G网络路由场景

当前5G及未来网络承载着具有多样化服务质量要求的异构流量,实时路由决策既复杂又关键。传统方法如人工干预的启发式策略、单一中心化强化学习或同步多学习者更新,存在可扩展性差和延迟节点问题。本文提出异步多智能体强化学习(AMARL)框架:每个服务配置一个独立的PPO智能体,在并行规划路径的同时,向共享全局资源环境提交资源变动量。这种基于状态的协调机制确保了各服务的可行性,并支持针对服务特性的目标优化。我们在类O-RAN网络仿真中使用蒙特利尔市近实时交通数据进行评估,对比单智能体PPO基线。AMARL在服务接受率(GoS)和端到端延迟上表现相当,同时显著降低训练耗时,并对需求波动更具鲁棒性。结果表明,异步、服务专用的智能体为分布式路由提供了可扩展且实用的解决方案,其应用潜力超出O-RAN领域。

原文摘要 · Abstract (English)

Networks in the current 5G and beyond systems increasingly carry heterogeneous traffic with diverse quality-of-service constraints, making real-time routing decisions both complex and time-critical. A common approach, such as a heuristic with human intervention or training a single centralized RL policy or synchronizing updates across multiple learners, struggles with scalability and straggler effects. We address this by proposing an asynchronous multi-agent reinforcement learning (AMARL) framework in which independent PPO agents, one per service, plan routes in parallel and commit resource deltas to a shared global resource environment. This coordination by state preserves feasibility across services and enables specialization for service-specific objectives. We evaluate the method on an O-RAN like network simulation using nearly real-time traffic data from the city of Montreal. We compared against a single-agent PPO baseline. AMARL achieves a similar Grade of Service (acceptance rate) (GoS) and end-to-end latency, with reduced training wall-clock time and improved robustness to demand shifts. These results suggest that asynchronous, service-specialized agents provide a scalable and practical approach to distributed routing, with applicability extending beyond the O-RAN domain.

5G路由强化学习多智能体异步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。