用强化学习动态调整5G开放基站的模型更新,省成本还保精度。
ADORN: Adaptive Drift handling for Open RAN using Reinforcement Learning

- 设计Q-learning智能体,按需决定是否重训练模型。
- 实测比盲目和随机更新少30%以上重训练开销。
- 多专家LSTM集成抗遗忘,适配复杂网络流量变化。
开放无线接入网(O-RAN)中动态流量变化导致模型漂移,降低人工智能/机器学习(AI/ML)模型性能。传统重训方法虽能保持预测精度,但计算开销大,且可能违反服务等级协议(SLA)。本文提出基于Q-learning的自适应重训方法,将重训决策建模为马尔可夫决策过程(MDP),让强化学习(RL)代理学习在预测精度与重训成本间权衡的策略。该方法采用多专家长短期记忆(LSTM)集成,缓解灾难性遗忘,提升在多样化流量条件下的鲁棒性。实验结果表明,相比贪婪和随机基线,该方法显著降低重训开销,同时将系统性能控制在预设范围内。
原文摘要 · Abstract (English)
Dynamic traffic variations in Open Radio Access Networks (O-RAN) lead to drift, which degrades the performance of Artificial Intelligence/Machine Learning (AI/ML) models. Traditional retraining approaches maintain forecasting accuracy but incur high computational cost and may lead to violations of Service Level Agreements (SLAs). This work proposes a Q-learning-based adaptive retraining approach that formulates the retraining decision as a Markov Decision Process (MDP), where a Reinforcement Learning (RL) agent learns a policy that balances forecasting accuracy and retraining cost. The proposed approach incorporates a multi-expert Long Short-Term Memory (LSTM) ensemble to mitigate catastrophic forgetting and improve robustness across diverse traffic conditions. Experimental results show that the proposed approach effectively reduces retraining overhead compared to greedy and random baselines, while maintaining system performance within predefined limits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。