用多智能体强化学习解决卫星通信中因延迟导致的信道信息失效问题
Multi-Agent Reinforcement Learning Counteracts Delayed CSI in Multi-Satellite Systems
- 提出双阶段近端策略优化算法,分步优化单星与多星协作下的传输速率
- 在信道状态信息滞后场景下,使用户总速率提升约18.7%且保持稳定
- 适合研究卫星通信、智能资源调度及强化学习应用的科研人员参考
将卫星通信网络与下一代(NG)技术融合是实现全球连接的有前景路径。然而,服务质量高度依赖于准确的信道状态信息(CSI)。由于地面用户与卫星间存在高传播延迟,导致卫星侧获得的信道估计严重滞后。本文研究多个卫星作为分布式基站向移动地面用户进行下行传输的问题。提出一种多智能体强化学习(MARL)算法,旨在最大化用户总速率的同时应对过时的CSI。设计了一种新颖的双层优化框架,称为双阶段近端策略优化(DS-PPO),以解决大规模连续动作空间及独立非同分布(non-IID)环境带来的挑战。具体而言,第一阶段针对单个卫星最大化总速率,第二阶段则实现所有卫星协同形成分布式多天线基站时的全局最优。数值结果表明,DS-PPO对CSI不完善具有强鲁棒性,并带来约18.7%的速率增益。此外,本文还提供了DS-PPO的收敛性分析与计算复杂度评估。
原文摘要 · Abstract (English)
The integration of satellite communication networks with next-generation (NG) technologies is a promising approach towards global connectivity. However, the quality of services is highly dependant on the availability of accurate channel state information (CSI). Channel estimation in satellite communications is challenging due to the high propagation delay between terrestrial users and satellites, which results in outdated CSI observations on the satellite side. In this paper, we study the downlink transmission of multiple satellites acting as distributed base stations (BS) to mobile terrestrial users. We propose a multi-agent reinforcement learning (MARL) algorithm which aims for maximising the sum-rate of the users, while coping with the outdated CSI. We design a novel bi-level optimisation, procedure themes as dual stage proximal policy optimisation (DS-PPO), for tackling the problem of large continuous action spaces as well as of independent and non-identically distributed (non-IID) environments in MARL. Specifically, the first stage of DS-PPO maximises the sum-rate for an individual satellite and the second stage maximises the sum-rate when all the satellites cooperate to form a distributed multi-antenna BS. Our numerical results demonstrate the robustness of DS-PPO to CSI imperfections as well as the sum-rate improvement attached by the use of DS-PPO. In addition, we provide the convergence analysis for the DS-PPO along with the computational complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。