用智能表面与无人机协同优化通信系统,提升频谱效率。
Sum Rate Maximization in STAR-RIS-UAV-Assisted Networks: A CA-DDPG Approach for Joint Optimization
- 基于强化学习联合优化波束成形、相位偏移和无人机位置。
- 新算法在仿真中使系统总速率显著提升,优于其他方法。
- 适合研究无线网络优化与智能反射面应用的学者参考。
随着可编程材料的发展,可重构智能表面(RIS)已成为未来无线通信的关键技术。同时具备发射与反射能力的可 simultaneously transmitting and reflecting RIS(STAR-RIS)能够全面控制信号,拓展应用场景。本文引入无人机(UAV)以进一步提升系统灵活性,并提出一种针对STAR-RIS-UAV辅助无线通信系统的频谱效率优化设计。我们提出一种深度强化学习(DRL)算法,通过与环境持续交互,迭代优化波束成形、相位偏移和无人机位置,以最大化系统总速率。为增强确定性策略下的探索能力,引入随机扰动因子。随着探索能力提升,状态-动作值函数的精确评估变得关键。因此,在深度确定性策略梯度(DDPG)基础上,提出卷积增强型深度确定性策略梯度(CA-DDPG)算法,平衡探索与评估,提升系统总速率。仿真结果表明,该算法能有效与环境交互,优化波束成形矩阵、相位偏移矩阵及无人机位置,提升系统容量,性能优于其他算法。
原文摘要 · Abstract (English)
With the rapid advances in programmable materials, reconfigurable intelligent surfaces (RIS) have become a pivotal technology for future wireless communications. The simultaneous transmitting and reflecting reconfigurable intelligent surfaces (STAR-RIS) can both transmit and reflect signals, enabling comprehensive signal control and expanding application scenarios. This paper introduces an unmanned aerial vehicle (UAV) to further enhance system flexibility and proposes an optimization design for the spectrum efficiency of the STAR-RIS-UAV-assisted wireless communication system. We present a deep reinforcement learning (DRL) algorithm capable of iteratively optimizing beamforming, phase shifts, and UAV positioning to maximize the system's sum rate through continuous interactions with the environment. To improve exploration in deterministic policies, we introduce a stochastic perturbation factor, which enhances exploration capabilities. As exploration is strengthened, the algorithm's ability to accurately evaluate the state-action value function becomes critical. Thus, based on the deep deterministic policy gradient (DDPG) algorithm, we propose a convolution-augmented deep deterministic policy gradient (CA-DDPG) algorithm that balances exploration and evaluation to improve the system's sum rate. The simulation results demonstrate that the CA-DDPG algorithm effectively interacts with the environment, optimizing the beamforming matrix, phase shift matrix, and UAV location, thereby improving system capacity and achieving better performance than other algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。