arXiv:2505.05956eess.SPcs.LG2025-05中稿 · Presentation at IE…被引 1

用深度强化学习优化毫米波多用户波束,提升通信稳定性和吞吐量。

Multi-User Beamforming with Deep Reinforcement Learning in Sensing-Aided Communication

  • 通过深度强化学习动态分配多波束与单波束,适应用户角度估计需求。
  • 相比传统波束扫描和启发式方法,吞吐量显著提升,且在不同速度下表现稳健。
  • 仅依赖感知回波,无需用户反馈或状态先验信息,适合实际部署。

毫米波通信中移动用户易因波束漂移导致波束失败。感知技术可通过及时波束更新和低开销缓解此问题,且无需用户反馈。本文研究通过动态管理分配给移动用户的波束来优化感知辅助通信。提出一种多波束方案:对需更新发射角(AoD)估计的用户分配多个波束,对已满足精度要求的用户分配单个波束。设计了一种基于深度强化学习(DRL)的波束分配策略,仅依赖感知回波。作为对比,还提出一种基于近似克拉美-罗界(CRLB)的启发式AoD分配方法。两种方法均无需用户反馈或状态演化先验信息。结果表明,该DRL方法在吞吐量上显著优于传统波束扫描法和AoD方法,且对不同用户速度具有鲁棒性。

原文摘要 · Abstract (English)

Mobile users are prone to experience beam failure due to beam drifting in millimeter wave (mmWave) communications. Sensing can help alleviate beam drifting with timely beam changes and low overhead since it does not need user feedback. This work studies the problem of optimizing sensing-aided communication by dynamically managing beams allocated to mobile users. A multi-beam scheme is introduced, which allocates multiple beams to the users that need an update on the angle of departure (AoD) estimates and a single beam to the users that have satisfied AoD estimation precision. A deep reinforcement learning (DRL) assisted method is developed to optimize the beam allocation policy, relying only upon the sensing echoes. For comparison, a heuristic AoD-based method using approximated Cramér-Rao lower bound (CRLB) for allocation is also presented. Both methods require neither user feedback nor prior state evolution information. Results show that the DRL-assisted method achieves a considerable gain in throughput than the conventional beam sweeping method and the AoD-based method, and it is robust to different user speeds.

毫米波通信强化学习波束成形感知辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。