arXiv:2505.08032cs.NIcs.AI2025-05被引 3

用在线强化学习提升6G波束切换的稳定性和效率

Online Learning-based Adaptive Beam Switching for 6G Networks: Enhancing Efficiency and Resilience

  • 引入包含阻塞历史的状态表示和以稳定性为核心的奖励函数
  • 在100用户场景下使链路稳定性提升43%,接近经典基线表现
  • 适合对实时性与可靠性要求高的军事及商用6G网络应用

自适应波束切换对于任务关键型军用和商用6G网络至关重要,但面临高频载波、用户移动性和频繁阻塞等挑战。现有机器学习方法多追求瞬时吞吐量最大化,导致策略不稳定且信令开销高。本文提出一种基于在线深度强化学习(DRL)的框架,通过增强状态表示(包含阻塞历史)和稳定性导向的奖励函数,使代理更关注长期链路质量而非短期收益。在使用Sionna库构建的100用户复杂场景中验证,该框架吞吐量与反应式多臂赌博机(MAB)基线相当。相比原始DRL方法,链路稳定性提升约43%,在保持高数据率的同时实现与MAB相当的运行可靠性。研究表明,将优化目标转向操作稳定性,可使DRL在下一代任务关键型网络中提供高效、可靠且实时的波束管理方案。

原文摘要 · Abstract (English)

Adaptive beam switching is essential for mission-critical military and commercial 6G networks but faces major challenges from high carrier frequencies, user mobility, and frequent blockages. While existing machine learning (ML) solutions often focus on maximizing instantaneous throughput, this can lead to unstable policies with high signaling overhead. This paper presents an online Deep Reinforcement Learning (DRL) framework designed to learn an operationally stable policy. By equipping the DRL agent with an enhanced state representation that includes blockage history, and a stability-centric reward function, we enable it to prioritize long-term link quality over transient gains. Validated in a challenging 100-user scenario using the Sionna library, our agent achieves throughput comparable to a reactive Multi-Armed Bandit (MAB) baseline. Specifically, our proposed framework improves link stability by approximately 43% compared to a vanilla DRL approach, achieving operational reliability competitive with MAB while maintaining high data rates. This work demonstrates that by reframing the optimization goal towards operational stability, DRL can deliver efficient, reliable, and real-time beam management solutions for next-generation mission-critical networks.

6G网络波束切换强化学习稳定性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。