用双智能体强化学习优化空中智能表面的通信效率与公平性。
Dual Actor DDPG for Airborne STAR-RIS Assisted Communications
- 设计双演员网络处理无人机轨迹、波束成形和反射系数联合优化。
- 相比传统方法,累积奖励提升24%以上,用户服务质量拒绝率降低41%。
- 适合研究无线网络优化、智能表面系统或强化学习应用的科研人员。
本研究突破了空中同时发射与反射可重构智能表面(Aerial-STAR)中发射与反射系数独立的传统假设,提出一种新型多用户下行链路通信系统,采用耦合相位移模型的无人机搭载STAR-RIS。核心贡献在于联合优化无人机轨迹、基站主动波束成形向量及无源RIS的发射/反射系数(TRC),同时考虑无人机能量约束。将TRC建模为离散与连续动作的组合,提出新型双演员深度确定性策略梯度(DA-DDPG)算法,利用两个独立演员网络处理高维混合动作空间。设计基于调和均值指数(HFI)的奖励函数以保障用户间通信公平性。仿真表明,所提算法在累积奖励上较传统DDPG和DQN分别提升24%和97%;三维轨迹优化相较二维或高度优化提升28%通信效率;HFI奖励函数使服务质量拒绝率降低41%。移动式Aerial-STAR系统性能优于固定部署,耦合相位STAR-RIS表现优于分离发射/反射RIS与传统RIS架构。结果验证了Aerial-STAR系统的潜力及所提方法的有效性。
原文摘要 · Abstract (English)
This study departs from the prevailing assumption of independent Transmission and Reflection Coefficients (TRC) in Airborne Simultaneous Transmit and Reflect Reconfigurable Intelligent Surface (STAR-RIS) research. Instead, we explore a novel multi-user downlink communication system that leverages a UAV-mounted STAR-RIS (Aerial-STAR) incorporating a coupled TRC phase shift model. Our key contributions include the joint optimization of UAV trajectory, active beamforming vectors at the base station, and passive RIS TRCs to enhance communication efficiency, while considering UAV energy constraints. We design the TRC as a combination of discrete and continuous actions, and propose a novel Dual Actor Deep Deterministic Policy Gradient (DA-DDPG) algorithm. The algorithm relies on two separate actor networks for high-dimensional hybrid action space. We also propose a novel harmonic mean index (HFI)-based reward function to ensure communication fairness amongst users. For comprehensive analysis, we study the impact of RIS size on UAV aerodynamics showing that it increases drag and energy demand. Simulation results demonstrate that the proposed DA-DDPG algorithm outperforms conventional DDPG and DQN-based solutions by 24% and 97%, respectively, in accumulated reward. Three-dimensional UAV trajectory optimization achieves 28% higher communication efficiency compared to two-dimensional and altitude optimization. The HFI based reward function provides 41% lower QoS denial rates as compared to other benchmarks. The mobile Aerial-STAR system shows superior performance over fixed deployed counterparts, with the coupled phase STAR-RIS outperforming dual Transmit/Reflect RIS and conventional RIS setups. These findings highlight the potential of Aerial-STAR systems and the effectiveness of our proposed DA-DDPG approach in optimizing their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。