arXiv:2507.06602cs.LG2025-07被引 4

用强化学习提升5G基站调度,让模型在不同网络环境里都表现稳定。

Generalization in Reinforcement Learning for Radio Access Networks

  • 通过图神经网络和部分观测重建,融合动态与静态网络信息。
  • 在五组5G基准测试中,吞吐量平均提升10%,高速移动下超20%。
  • 适合希望部署统一智能调度系统的运营商或6G研究者。

现代无线接入网(RAN)运行在高度动态和异构的环境中,传统手工调参的资源管理算法常表现不佳。尽管强化学习(RL)在受限场景下可超越启发式方法,但部署多样性与不可预测的无线条件带来显著泛化挑战。数据驱动策略常在训练条件下过拟合,导致新场景性能下降。为此,本文提出一种以泛化为核心的RAN控制强化学习框架:(i)从不完整、噪声干扰的观测中鲁棒地重构动态状态,同时利用图表示编码静态与半静态信息,如无线节点、小区属性及其拓扑;(ii)采用领域随机化扩展训练分布;(iii)在多智能体并行生成数据的同时,采用云兼容架构集中训练,符合O-RAN原则。尽管泛化增加计算与数据管理复杂度,分布式设计通过跨多样化网络条件扩展数据收集与训练加以缓解。应用于五个5G下行链路自适应基准测试,所提策略在满缓冲区MIMO/mMIMO场景下,平均吞吐量与频谱效率较基线(OLLA,BLER目标10%)提升约10%,高移动性下超过20%。在满缓冲流量下媲美专用强化学习模型,在eMBB与混合流量基准上分别实现4倍与2倍性能增益。九小区部署中,图注意力网络(GAT)相较多层感知机(MLP)提升30%吞吐量。结合可扩展架构,该方案为基于单一通用强化学习代理实现6G原生智能无线接入网提供了可行路径。

原文摘要 · Abstract (English)

Modern RAN operate in highly dynamic and heterogeneous environments, where hand-tuned, rule-based RRM algorithms often underperform. While RL can surpass such heuristics in constrained settings, the diversity of deployments and unpredictable radio conditions introduce major generalization challenges. Data-driven policies frequently overfit to training conditions, degrading performance in unseen scenarios. To address this, we propose a generalization-centered RL framework for RAN control that: (i) robustly reconstructs dynamically varying states from partial and noisy observations, while encoding static and semi-static information, such as radio nodes, cell attributes, and their topology, through graph representations; (ii) applies domain randomization to broaden the training distribution; and (iii) distributes data generation across multiple actors while centralizing training in a cloud-compatible architecture aligned with O-RAN principles. Although generalization increases computational and data-management complexity, our distributed design mitigates this by scaling data collection and training across diverse network conditions. Applied to downlink link adaptation in five 5G benchmarks, our policy improves average throughput and spectral efficiency by ~10% over an OLLA baseline (10% BLER target) in full-buffer MIMO/mMIMO and by >20% under high mobility. It matches specialized RL in full-buffer traffic and achieves up to 4- and 2-fold gains in eMBB and mixed-traffic benchmarks, respectively. In nine-cell deployments, GAT models offer 30% higher throughput over MLP baselines. These results, combined with our scalable architecture, offer a path toward AI-native 6G RAN using a single, generalizable RL agent.

强化学习5G优化无线网络图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。