用强化学习动态追踪5G/6G中移动用户的最优波束,提升连接稳定性。
Meta-Learning Multi-armed Bandits for Beam Tracking in 5G and 6G Networks
- 将波束选择建模为部分可观测马尔可夫决策过程,基于信念状态搜索最优波束。
- 在复杂反射与遮挡环境下,性能优于传统方法数个数量级。
- 适用于未知轨迹或环境突变场景,适合通信系统优化研究者。
具备大量天线元素的波束成形阵列可提升下一代5G和6G网络的数据速率。当前实践中,模拟波束成形使用预配置的波束码本,每个波束指向特定方向,波束管理功能持续为移动用户设备(UE)选择最优波束。然而,大码本及波束反射或遮挡效应使最优波束选择极具挑战。不同于以往通过监督学习训练分类器预测下一最佳波束的方法,本文将问题建模为部分可观测马尔可夫决策过程(POMDP),将码本本身视为环境。每一步时间,根据不可观测最优波束的信念状态及先前探测波束,选择候选波束。这将波束选择问题转化为在线搜索过程,以定位移动中的最优波束。相较于以往工作,该方法能处理新路径或物理环境变化,性能提升达数量级。
原文摘要 · Abstract (English)
Beamforming-capable antenna arrays with many elements enable higher data rates in next generation 5G and 6G networks. In current practice, analog beamforming uses a codebook of pre-configured beams with each of them radiating towards a specific direction, and a beam management function continuously selects \textit{optimal} beams for moving user equipments (UEs). However, large codebooks and effects caused by reflections or blockages of beams make an optimal beam selection challenging. In contrast to previous work and standardization efforts that opt for supervised learning to train classifiers to predict the next best beam based on previously selected beams we formulate the problem as a partially observable Markov decision process (POMDP) and model the environment as the codebook itself. At each time step, we select a candidate beam conditioned on the belief state of the unobservable optimal beam and previously probed beams. This frames the beam selection problem as an online search procedure that locates the moving optimal beam. In contrast to previous work, our method handles new or unforeseen trajectories and changes in the physical environment, and outperforms previous work by orders of magnitude.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。