针对6G超低时延通信中的突发流量,提出混合强化学习调度框架。
A Hybrid Reinforcement Learning Framework for Hard Latency Constrained Resource Scheduling
- 融合历史策略与专家知识构建新策略,提升资源调度效率。
- 在硬时延约束下,有效吞吐量提升显著,收敛速度更快。
- 适合研究6G URLLC中高可靠低时延场景的系统设计者。
在未来的6G时代,扩展现实(XR)被视为超可靠低时延通信(URLLC)的新兴应用,具有新的流量特征和更严格的时延要求。除了XR中的准周期性流量外,某些真实低时延通信场景中的大帧突发流量(随机到达、高突发性)已成为导致网络拥塞甚至崩溃的主要原因,且目前尚无高效算法应对具有硬时延约束的突发流量资源调度问题。本文提出一种新型混合强化学习资源调度框架(HRL-RSHLC),通过复用其他相似环境中学到的旧策略以及基于领域知识(DK)构造的专家策略,提升性能。将策略复用概率与新策略联合优化建模为马尔可夫决策过程(MDP),以最大化用户在硬时延约束下的有效吞吐量(HLC-ET)。理论证明,该框架可从任意初始点收敛至KKT点。仿真结果表明,相较于基线算法,HRL-RSHLC在硬时延约束下表现出更优性能,且收敛速度更快。
原文摘要 · Abstract (English)
In the forthcoming 6G era, extend reality (XR) has been regarded as an emerging application for ultra-reliable and low latency communications (URLLC) with new traffic characteristics and more stringent requirements. In addition to the quasi-periodical traffic in XR, burst traffic with both large frame size and random arrivals in some real world low latency communication scenarios has become the leading cause of network congestion or even collapse, and there still lacks an efficient algorithm for the resource scheduling problem under burst traffic with hard latency constraints. We propose a novel hybrid reinforcement learning framework for resource scheduling with hard latency constraints (HRL-RSHLC), which reuses polices from both old policies learned under other similar environments and domain-knowledge-based (DK) policies constructed using expert knowledge to improve the performance. The joint optimization of the policy reuse probabilities and new policy is formulated as an Markov Decision Problem (MDP), which maximizes the hard-latency constrained effective throughput (HLC-ET) of users. We prove that the proposed HRL-RSHLC can converge to KKT points with an arbitrary initial point. Simulations show that HRL-RSHLC can achieve superior performance with faster convergence speed compared to baseline algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。