用强化学习动态调度用户优先级和功率,提升6G网络能效与服务质量。
Joint User Priority and Power Scheduling for QoS-Aware WMMSE Precoding: A Constrained-Actor Attentive-Critic Approach
- 设计基于约束强化学习的注意力策略网络,动态分配用户优先级和发射功率。
- 在仿真中实现更高能效与95%以上用户服务质量满足率。
- 适合研究6G资源调度、智能无线优化的学者与工程师参考。
6G无线网络需在保障高能效的同时满足多样化的服务质量(QoS)需求。传统的加权最小均方误差(WMMSE)预编码虽能提升系统性能,但固定用户优先级和发射功率缺乏灵活性,难以适应用户特定的QoS要求及时变信道。为此,本文提出一种新型约束强化学习算法——受限-演员注意力评价者(CAAC),通过策略网络动态分配用户优先级与功率。具体而言,CAAC结合受限随机逐次凸逼近(CSSCA)优化策略,更有效地处理能效目标并满足随机非凸QoS约束,优于传统及现有约束强化学习方法。此外,CAAC采用轻量级注意力增强的Q网络评估策略更新,无需环境模型先验知识,兼具更强表征能力与学习效率。仿真结果表明,相较于基线方法,CAAC在能效和QoS满足率方面均有显著提升。
原文摘要 · Abstract (English)
6G wireless networks are expected to support diverse quality-of-service (QoS) demands while maintaining high energy efficiency. Weighted Minimum Mean Square Error (WMMSE) precoding with fixed user priorities and transmit power is widely recognized for enhancing overall system performance but lacks flexibility to adapt to user-specific QoS requirements and time-varying channel conditions. To address this, we propose a novel constrained reinforcement learning (CRL) algorithm, Constrained-Actor Attentive-Critic (CAAC), which uses a policy network to dynamically allocate user priorities and power for WMMSE precoding. Specifically, CAAC integrates a Constrained Stochastic Successive Convex Approximation (CSSCA) method to optimize the policy, enabling more effective handling of energy efficiency goals and satisfaction of stochastic non-convex QoS constraints compared to traditional and existing CRL methods. Moreover, CAAC employs lightweight attention-enhanced Q-networks to evaluate policy updates without prior environment model knowledge. The network architecture not only enhances representational capacity but also boosts learning efficiency. Simulation results show that CAAC outperforms baselines in both energy efficiency and QoS satisfaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。