用深度强化学习提升5G高速移动下的链路自适应,吞吐量最高提升92%。
LOLLA: Deep Reinforcement Learning for Closed-Loop Link Adaptation Towards a GPU-Accelerated AI-RAN

- 用深度强化学习替代传统链路自适应,根据丰富信道数据动态调整信号质量阈值。
- 在400Hz多普勒频移下,吞吐量比传统方法高15%至92%,且可靠性可调。
- 首个基于GPU的5G端到端闭环AI控制应用,支持多用户并发和真实部署平滑迁移。
外环链路自适应(OLLA)广泛应用于5G NR以追踪信道变化,但在高速移动和快速时变信道下,其依赖的一阶单比特反馈会显著降低性能。本文提出LOLLA(学习型外环链路自适应),一种深度强化学习框架,用可学习的连续信噪比偏置替代传统阶梯式更新,该偏置基于物理层与媒体接入层的丰富遥测数据,且不破坏3GPP兼容的调制编码策略选择,并可形式化包含传统更新规则。采用拉格朗日约束块错误率(BLER)的近端策略优化(PPO)策略,在无需人工调节惩罚参数的情况下,实现1%至15%可调的可靠目标。该框架首次实现在GPU加速5G NR栈上的闭环AI原生控制dApp,端到端控制延迟低于500微秒。在3GPP TDL信道模型下评估显示,相比传统OLLA,吞吐量提升15%至92%(最高达400 Hz多普勒频率),且在所有可靠性目标下均严格优于传统方案,达到帕累托最优。所学策略对未见过的信道模型具备泛化能力,并可在共享资源调度下支持8个并发用户。上行链路中,基站直接观测解码结果,实现仿真到部署的无缝一致性。
原文摘要 · Abstract (English)
Outer-loop link adaptation (OLLA) is widely deployed in 5G NR to track channel variations, yet its reliance on first-order, single-bit feedback degrades performance significantly under high-mobility and fast-varying channels. This paper presents LOLLA (Learned Outer-Loop Link Adaptation), a deep reinforcement learning framework that replaces the conventional OLLA staircase with a learned, continuous SINR offset conditioned on rich PHY/MAC telemetry inaccessible to OLLA. The offset modulates the SINR-to-MCS lookup table, preserving 3GPP-compliant MCS selection and provably subsuming the conventional OLLA update rule. A Proximal Policy Optimization (PPO) policy trained under a Lagrangian block error rate (BLER) constraint automatically enforces tunable reliability targets from 1% to 15% without manual penalty calibration. The framework is realized as the first closed-loop AI-native control dApp on a GPU-accelerated 5G NR stack, achieving end-to-end control latencies under 500 microseconds. Evaluations under 3GPP TDL channel models demonstrate 15% to 92% throughput gains over OLLA across Doppler frequencies up to 400 Hz, while attaining a Pareto frontier that strictly dominates OLLA across all evaluated reliability targets. The learned policy generalizes to unseen channel models and scales to eight concurrent UEs under shared-resource scheduling. In the uplink formulation, the gNB directly observes decoding outcomes, enabling simulation-to-deployment parity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。