arXiv:2503.09391eess.SYcs.ET2025-03

针对非平稳XR流量,提出上下文感知强化学习实现节能调度。

Context-aware Constrained Reinforcement Learning Based Energy-Efficient Power Scheduling for Non-stationary XR Data Traffic

  • 用上下文推理模块捕捉动态网络状态,重塑稀疏奖励为密集反馈。
  • 在仿真中降低32%能耗,同时满足98%的丢包率约束。
  • 适合5G/6G时代高延迟敏感场景的智能资源调度研究者。

在XR下行传输中,能源高效功率调度(EEPS)对于在硬性时延约束下传输大数据包并节约能源资源至关重要。传统受限强化学习(CRL)算法虽有潜力,但在非凸随机约束、非平稳数据流量及稀疏延迟丢包反馈方面仍面临挑战。本文将XR中的EEPS建模为具有随非平稳流量变化的转移函数的动态参数受限马尔可夫决策过程(DP-CMDP),并提出一种上下文感知受限强化学习(CACRL)算法,包含上下文推理(CI)模块和CRL模块。CI模块通过训练编码器与多个潜在网络,刻画当前转移函数,并根据上下文重置丢包奖励,将原DP-CMDP转化为具即时密集奖励的一般CMDP。CRL模块使用策略网络在该CMDP下做出调度决策,并采用更适合非凸随机约束的受限随机连续凸近似(CSSCA)方法优化策略。理论分析揭示了算法内在机理,大量仿真表明其在节能与满足丢包约束方面均优于先进基线。

原文摘要 · Abstract (English)

In XR downlink transmission, energy-efficient power scheduling (EEPS) is essential for conserving power resource while delivering large data packets within hard-latency constraints. Traditional constrained reinforcement learning (CRL) algorithms show promise in EEPS but still struggle with non-convex stochastic constraints, non-stationary data traffic, and sparse delayed packet dropout feedback (rewards) in XR. To overcome these challenges, this paper models the EEPS in XR as a dynamic parameter-constrained Markov decision process (DP-CMDP) with a varying transition function linked to the non-stationary data traffic and solves it by a proposed context-aware constrained reinforcement learning (CACRL) algorithm, which consists of a context inference (CI) module and a CRL module. The CI module trains an encoder and multiple potential networks to characterize the current transition function and reshape the packet dropout rewards according to the context, transforming the original DP-CMDP into a general CMDP with immediate dense rewards. The CRL module employs a policy network to make EEPS decisions under this CMDP and optimizes the policy using a constrained stochastic successive convex approximation (CSSCA) method, which is better suited for non-convex stochastic constraints. Finally, theoretical analyses provide deep insights into the CADAC algorithm, while extensive simulations demonstrate that it outperforms advanced baselines in both power conservation and satisfying packet dropout constraints.

强化学习节能调度XR通信非平稳流量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。