用深度强化学习优化低空经济中的通感一体化系统
Integrated Sensing and Communications for Low-Altitude Economy: A Deep Reinforcement Learning Approach
- 通过深度强化学习联合优化基站波束成形与无人机轨迹
- 在满足感知信噪比和飞行安全约束下,通信速率提升显著
- 适合低空交通管理、无人机调度等实际应用场景
本文研究面向低空经济(LAE)的通感一体化(ISAC)系统,地面基站(GBS)为授权无人机(UAV)提供通信与导航服务,同时监测低空空域以发现非法移动目标。通过联合优化GBS的波束成形与无人机轨迹,在满足平均感知信噪比要求、飞行任务、避撞约束及最大发射功率限制的前提下,最大化给定飞行周期内的通信总速率。该问题具有序列决策特性,因此建模为特定马尔可夫决策过程(MDP),称为“回合任务”。基于此,提出面向低空经济的新型ISAC方案DeepLSC,利用深度强化学习技术。设计了合理的奖励函数与带约束的噪声探索策略以满足各类约束。为提升学习效率,引入分层经验回放机制,利用每回合内所有经验联合训练神经网络;为进一步加速收敛,提出对称经验增强机制,通过同时置换所有变量索引丰富经验集。仿真结果表明,相比基准方法,DeepLSC在满足预设约束的同时,实现更高通信总速率、更快收敛速度,并具备更强鲁棒性。
原文摘要 · Abstract (English)
This paper studies an integrated sensing and communications (ISAC) system for low-altitude economy (LAE), where a ground base station (GBS) provides communication and navigation services for authorized unmanned aerial vehicles (UAVs), while sensing the low-altitude airspace to monitor the unauthorized mobile target. The expected communication sum-rate over a given flight period is maximized by jointly optimizing the beamforming at the GBS and UAVs' trajectories, subject to the constraints on the average signal-to-noise ratio requirement for sensing, the flight mission and collision avoidance of UAVs, as well as the maximum transmit power at the GBS. Typically, this is a sequential decision-making problem with the given flight mission. Thus, we transform it to a specific Markov decision process (MDP) model called episode task. Based on this modeling, we propose a novel LAE-oriented ISAC scheme, referred to as Deep LAE-ISAC (DeepLSC), by leveraging the deep reinforcement learning (DRL) technique. In DeepLSC, a reward function and a new action selection policy termed constrained noise-exploration policy are judiciously designed to fulfill various constraints. To enable efficient learning in episode tasks, we develop a hierarchical experience replay mechanism, where the gist is to employ all experiences generated within each episode to jointly train the neural network. Besides, to enhance the convergence speed of DeepLSC, a symmetric experience augmentation mechanism, which simultaneously permutes the indexes of all variables to enrich available experience sets, is proposed. Simulation results demonstrate that compared with benchmarks, DeepLSC yields a higher sum-rate while meeting the preset constraints, achieves faster convergence, and is more robust against different settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。