用多智能体强化学习实现低延迟服务的动态调度与路径规划
A Flexible Multi-Agent Deep Reinforcement Learning Framework for Dynamic Routing and Scheduling of Latency-Critical Services
- 采用集中路由+分布式调度架构,结合MADDPG算法动态分配路径和传输时机
- 在真实网络场景中实现98.7%的即时送达率,显著优于传统优化方法
- 支持融合深度强化学习与规则策略,适合工业自动化等高实时性场景
在动态异构网络中及时传递延迟敏感信息,对工业自动化、自动驾驶和增强现实等交互式应用日益关键。现有网络控制方案大多仅关注平均延迟,难以提供严格的端到端(E2E)峰值延迟保障。本文针对延迟约束下的最大吞吐量(DCMT)动态网络控制问题,提出一种新型多智能体深度强化学习(MA-DRL)框架。该框架采用集中式路由与分布式调度架构,基于多智能体深度确定性策略梯度(MADDPG)技术,利用关键网络领域知识设计高效策略。路由与调度智能体根据数据包剩余寿命动态分配路径并安排传输,以最大化按时送达率。框架具备通用性,可融合数据驱动的深度强化学习(DRL)代理与传统规则策略,在性能与学习复杂度间取得平衡。实验结果表明,该框架优于传统随机优化方法,并揭示了DRL代理与新规则策略在高效高保真控制中的协同作用。
原文摘要 · Abstract (English)
Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is increasingly essential for a range of interactive applications, such as industrial automation, self-driving vehicles, and augmented reality. However, most existing network control solutions target only average delay performance, falling short of providing strict End-to-End (E2E) peak latency guarantees. This paper addresses the challenge of reliably delivering packets within application-imposed deadlines by leveraging recent advancements in Multi-Agent Deep Reinforcement Learning (MA-DRL). After introducing the Delay-Constrained Maximum-Throughput (DCMT) dynamic network control problem, and highlighting the limitations of current solutions, we present a novel MA-DRL network control framework that leverages a centralized routing and distributed scheduling architecture. The proposed framework leverages critical networking domain knowledge for the design of effective MA-DRL strategies based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) technique, where centralized routing and distributed scheduling agents dynamically assign paths and schedule packet transmissions according to packet lifetimes, thereby maximizing on-time packet delivery. The generality of the proposed framework allows integrating both data-driven \blue{Deep Reinforcement Learning (DRL)} agents and traditional rule-based policies in order to strike the right balance between performance and learning complexity. Our results confirm the superiority of the proposed framework with respect to traditional stochastic optimization-based approaches and provide key insights into the role and interplay between data-driven DRL agents and new rule-based policies for both efficient and high-performance control of latency-critical services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。