对比三种离线强化学习算法,发现保守Q-learning在随机网络中更稳定可靠。
Selecting Offline Reinforcement Learning Algorithms for Stochastic Network Control
- 对比贝尔曼、序列和混合三类离线强化学习方法
- 保守Q-learning在多种随机性下表现最稳健,政策鲁棒性最高
- 序列方法在高回报轨迹充足时可超越贝尔曼方法,适合数据丰富的场景
离线强化学习(Offline RL)是下一代无线网络的有前途方案,因在线探索不安全且可复用大量运行数据。然而,现有方法在真实随机动态(如信道衰落、噪声、流量移动性)下的表现仍不明确。本文在公开的随机电信环境(mobile-env)中评估基于贝尔曼(保守Q-learning)、序列(决策变换器)和混合(评论家引导决策变换器)的离线RL方法。结果表明,保守Q-learning在不同随机性来源下均产生更稳健的策略,适合作为生命周期驱动的AI管理框架的默认选择。序列方法表现依然强劲,在高回报轨迹充足时甚至优于贝尔曼方法。这些发现为面向O-RAN和未来6G功能的AI驱动网络控制流程中的算法选择提供了实用指导。
原文摘要 · Abstract (English)
Offline Reinforcement Learning (RL) is a promising approach for next-generation wireless networks, where online exploration is unsafe and large amounts of operational data can be reused across the model lifecycle. However, the behavior of offline RL algorithms under genuinely stochastic dynamics -- inherent to wireless systems due to fading, noise, and traffic mobility -- remains insufficiently understood. We address this gap by evaluating Bellman-based (Conservative Q-Learning), sequence-based (Decision Transformers), and hybrid (Critic-Guided Decision Transformers) offline RL methods in an open-access stochastic telecom environment (mobile-env). Our results show that Conservative Q-Learning consistently produces more robust policies across different sources of stochasticity, making it a reliable default choice in lifecycle-driven AI management frameworks. Sequence-based methods remain competitive and can outperform Bellman-based approaches when sufficient high-return trajectories are available. These findings provide practical guidance for offline RL algorithm selection in AI-driven network control pipelines, such as O-RAN and future 6G functions, where robustness and data availability are key operational constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。