arXiv:2504.03804cs.LGcs.MA2025-04被引 9

用离线分布强化学习优化6G无线通信,提升安全与效率

Offline and Distributional Reinforcement Learning for Wireless Communications

  • 结合离线与分布强化学习,基于静态数据训练避免实时交互风险
  • 在无人机轨迹与资源管理中收敛更快,风险控制更优
  • 适合6G高可靠低时延场景,尤其需规避在线试错的系统

6G网络中异构海量连接的快速发展对智能解决方案提出新要求,需兼顾可扩展性、可靠性、隐私保护、超低延迟和有效控制。尽管人工智能与机器学习在此领域展现出潜力,但传统在线强化学习与深度强化学习方法在实时无线网络中受限:依赖环境实时交互,可能不可行、成本高或存在安全隐患,且难以应对真实无线应用中的固有不确定性。本文聚焦离线与分布强化学习两类先进方法,通过静态数据集训练并建模网络不确定性,克服上述挑战。提出一种融合离线与分布强化学习的新型框架,并在无人机轨迹优化与无线资源管理(RRM)案例中验证。结果表明,所提保守分位数回归(CQR)算法在收敛速度与风险管控方面优于传统强化学习方法。最后讨论了6G中应用这些技术的开放挑战与未来方向,为构建更安全高效的实时无线系统铺路。

原文摘要 · Abstract (English)

The rapid growth of heterogeneous and massive wireless connectivity in 6G networks demands intelligent solutions to ensure scalability, reliability, privacy, ultra-low latency, and effective control. Although artificial intelligence (AI) and machine learning (ML) have demonstrated their potential in this domain, traditional online reinforcement learning (RL) and deep RL methods face limitations in real-time wireless networks. For instance, these methods rely on online interaction with the environment, which might be unfeasible, costly, or unsafe. In addition, they cannot handle the inherent uncertainties in real-time wireless applications. We focus on offline and distributional RL, two advanced RL techniques that can overcome these challenges by training on static datasets and accounting for network uncertainties. We introduce a novel framework that combines offline and distributional RL for wireless communication applications. Through case studies on unmanned aerial vehicle (UAV) trajectory optimization and radio resource management (RRM), we demonstrate that our proposed Conservative Quantile Regression (CQR) algorithm outperforms conventional RL approaches regarding convergence speed and risk management. Finally, we discuss open challenges and potential future directions for applying these techniques in 6G networks, paving the way for safer and more efficient real-time wireless systems.

强化学习6G通信离线学习风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。