arXiv:2603.04353cs.NIcs.LG2026-03中稿 · publication in 202…

用强化学习优化实时应用的低延迟高成本效益传输

A Constrained RL Approach for Cost-Efficient Delivery of Latency-Sensitive Applications

  • 将网络调度建模为约束马尔可夫决策过程,用约束强化学习求解
  • 在严格单包延迟要求下仍保持高可靠及时交付,成本低于现有方法
  • 适合需要低延迟、低成本资源调度的实时服务场景

下一代网络致力于为实时交互服务提供性能保障,要求在满足应用严格截止期限的前提下实现高效且低成本的分组传输。已有大量研究采用随机优化技术设计动态路由与调度方案,以应对平均延迟约束;然而,当面临严格的单包延迟要求时,这些方法表现不足。本文将最小化成本的延迟约束网络控制问题建模为约束马尔可夫决策过程,并利用约束深度强化学习(CDRL)技术,在确保及时吞吐量高于目标可靠性水平的同时,有效降低总体资源分配成本。实验结果表明,所提方法能在现有基线失效的情况下仍保证及时分组交付,且相比其他以吞吐量最大化为目标的方法具有更低的成本。

原文摘要 · Abstract (English)

Next-generation networks aim to provide performance guarantees to real-time interactive services that require timely and cost-efficient packet delivery. In this context, the goal is to reliably deliver packets with strict deadlines imposed by the application while minimizing overall resource allocation cost. A large body of work has leveraged stochastic optimization techniques to design efficient dynamic routing and scheduling solutions under average delay constraints; however, these methods fall short when faced with strict per-packet delay requirements. We formulate the minimum-cost delay-constrained network control problem as a constrained Markov decision process and utilize constrained deep reinforcement learning (CDRL) techniques to effectively minimize total resource allocation cost while maintaining timely throughput above a target reliability level. Results indicate that the proposed CDRL-based solution can ensure timely packet delivery even when existing baselines fall short, and it achieves lower cost compared to other throughput-maximizing methods.

强化学习网络调度低延迟成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。