arXiv:2506.22793cs.NIcs.AI2025-06被引 1

用离线强化学习优化蜂窝网络切换参数,提升移动性鲁棒性。

Offline Reinforcement Learning for Mobility Robustness Optimization

  • 用决策变换器和保守Q学习从历史数据学最优切换偏移策略。
  • 在3500MHz NR网络上实现最高7%的性能提升,优于传统规则方法。
  • 同一数据集可适配多种目标函数,适合实际运维灵活调整。

本文重新研究移动性鲁棒性优化(MRO)算法,探索利用离线强化学习学习最优小区个体偏移(Cell Individual Offset)配置的可行性。此类方法基于收集的离线数据学习最优策略,无需额外探索。我们采用基于序列的决策变换器(Decision Transformers)和基于值函数的保守Q学习(Conservative Q-Learning),在与原始规则式MRO相同的目标奖励下进行训练。输入特征包括失败、乒乓切换及其他切换问题相关指标。在包含多样化用户业务类型、特定可调小区对的3500 MHz NR网络真实场景下评估表明,离线强化学习方法表现优于规则式MRO,最多提升7%。此外,同一数据集可用于训练多种目标函数,相较规则方法具备更高运营灵活性。

原文摘要 · Abstract (English)

In this work we revisit the Mobility Robustness Optimisation (MRO) algorithm and study the possibility of learning the optimal Cell Individual Offset tuning using offline Reinforcement Learning. Such methods make use of collected offline datasets to learn the optimal policy, without further exploration. We adapt and apply a sequence-based method called Decision Transformers as well as a value-based method called Conservative Q-Learning to learn the optimal policy for the same target reward as the vanilla rule-based MRO. The same input features related to failures, ping-pongs, and other handover issues are used. Evaluation for realistic New Radio networks with 3500 MHz carrier frequency on a traffic mix including diverse user service types and a specific tunable cell-pair shows that offline-RL methods outperform rule-based MRO, offering up to 7% improvement. Furthermore, offline-RL can be trained for diverse objective functions using the same available dataset, thus offering operational flexibility compared to rule-based methods.

强化学习蜂窝网络移动性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。