arXiv:2603.02745cs.ITcs.AI2026-03中稿 · the IEEE Internati…

用深度强化学习优化毫米波多天线网络的波束选择,提升用户吞吐量。

Enhancing User Throughput in Multi-panel mmWave Radio Access Networks for Beam-based MU-MIMO Using a DRL Method

  • 基于深度强化学习建模波束管理为马尔可夫决策过程,实时动态调整波束。
  • 实测显示吞吐量最高提升16%,端到端延迟降低3-7倍。
  • 适合关注毫米波通信、智能波束管理的系统设计与算法研究人员。

毫米波通信系统在采用混合波束成形的多用户多输入多输出(MU-MIMO)场景中,因动态波束选择与管理的高复杂性,面临用户吞吐量与延迟优化难题。本文提出一种基于深度强化学习(DRL)的方法,在实际多面板毫米波无线接入网络中提升用户吞吐量。所提框架将通信代理与环境的交互建模为马尔可夫决策过程(MDP),利用自适应波束管理策略,结合不同天线面板间波束的互相关性、测量的参考信号接收功率(RSRP)及波束使用统计信息,实现基于实时观测的波束选择优化。结果表明,该方法显著提升了频谱效率并降低了端到端延迟。数值实验显示,相比基线(传统波束管理),吞吐量最高提升16%,延迟降低3-7倍。

原文摘要 · Abstract (English)

Millimeter-wave (mmWave) communication systems, particularly those leveraging multi-user multiple-input and multiple-output (MU-MIMO) with hybrid beamforming, face challenges in optimizing user throughput and minimizing latency due to the high complexity of dynamic beam selection and management. This paper introduces a deep reinforcement learning (DRL) approach for enhancing user throughput in multi-panel mmWave radio access networks in a practical network setup. Our DRL-based formulation utilizes an adaptive beam management strategy that models the interaction between the communication agent and its environment as a Markov decision process (MDP), optimizing beam selection based on real-time observations. The proposed framework exploits spatial domain (SD) characteristics by incorporating the cross-correlation between the beams in different antenna panels, the measured reference signal received power (RSRP), and the beam usage statistics to dynamically adjust beamforming decisions. As a result, the spectral efficiency is improved and end-to-end latency is reduced. The numerical results demonstrate an increase in throughput of up to 16% and a reduction in latency by factors 3-7x compared to baseline (legacy beam management).

毫米波强化学习波束管理吞吐量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。