用离线强化学习优化无线链路自适应,不扰网也能达顶尖性能。
Offline Reinforcement Learning and Sequence Modeling for Downlink Link Adaptation
- 用离线强化学习替代在线训练,避免影响真实网络运行
- 在合理行为策略下,性能媲美前沿在线RL方法
- 适合通信系统研发与算法部署人员参考
链路自适应(LA)是现代无线通信系统中的关键功能,能根据时变、频变的信道条件动态调整传输速率。然而,用户移动性、快速衰落、信道质量信息不完整及测量数据老化等因素使LA建模极具挑战。为规避显式建模需求,近期研究提出在线强化学习(RL)作为规则算法的替代方案。但在线RL存在部署难题:在真实网络中训练可能损害实时性能。为此,本文探索离线强化学习作为学习LA策略的可行方案,可最大限度减少对网络运行的影响。提出三种基于批处理约束深度Q-learning、保守Q-learning和决策变压器的LA设计。实验表明,当数据由合适的策略采集时,离线RL算法性能可达到现有最优在线RL方法水平。
原文摘要 · Abstract (English)
Link adaptation (LA) is an essential function in modern wireless communication systems that dynamically adjusts the transmission rate of a communication link to match time- and frequency-varying radio link conditions. However, factors such as user mobility, fast fading, imperfect channel quality information, and aging of measurements make the modeling of LA challenging. To bypass the need for explicit modeling, recent research has introduced online reinforcement learning (RL) approaches as an alternative to the more commonly used rule-based algorithms. Yet, RL-based approaches face deployment challenges, as training in live networks can potentially degrade real-time performance. To address this challenge, this paper considers offline RL as a candidate to learn LA policies with minimal effects on the network operation. We propose three LA designs based on batch-constrained deep Q-learning, conservative Q-learning, and decision transformer. Our results show that offline RL algorithms can match the performance of state-of-the-art online RL methods when data is collected with a proper behavioral policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。