将李雅普诺夫稳定性机制融入强化学习,提升系统队列稳定性和长期性能。
A Lyapunov Drift-Plus-Penalty Method Tailored for Reinforcement Learning with Queue Stability
- 设计新型算法LDPTRLQ,融合李雅普诺夫漂移-惩罚与强化学习优势。
- 仿真显示该算法在多场景下优于基线方法,兼顾稳定与兼容性。
- 适合需要长期稳定运行的物联网智能决策场景使用。
随着物联网设备的普及,复杂优化问题需求日益增长。李雅普诺夫漂移-惩罚算法广泛用于保障队列稳定,已有研究初步探索其与强化学习(RL)的结合。本文针对强化学习应用,通过严谨的理论分析,提出一种适用于队列稳定的李雅普诺夫漂移-惩罚强化学习算法(LDPTRLQ)。不同于直接融合两个框架的现有方法,本算法在保证李雅普诺夫漂移-惩罚贪婪优化的同时,兼顾强化学习的长期视角,实现理论上的优越性。多种问题的仿真结果表明,LDPTRLQ显著优于基于李雅普诺夫漂移-惩罚与强化学习的基线方法,验证了理论推导的有效性,并在兼容性与稳定性方面表现更优。
原文摘要 · Abstract (English)
With the proliferation of Internet of Things (IoT) devices, the demand for addressing complex optimization challenges has intensified. The Lyapunov Drift-Plus-Penalty algorithm is a widely adopted approach for ensuring queue stability, and some research has preliminarily explored its integration with reinforcement learning (RL). In this paper, we investigate the adaptation of the Lyapunov Drift-Plus-Penalty algorithm for RL applications, deriving an effective method for combining Lyapunov Drift-Plus-Penalty with RL under a set of common and reasonable conditions through rigorous theoretical analysis. Unlike existing approaches that directly merge the two frameworks, our proposed algorithm, termed Lyapunov drift-plus-penalty method tailored for reinforcement learning with queue stability (LDPTRLQ) algorithm, offers theoretical superiority by effectively balancing the greedy optimization of Lyapunov Drift-Plus-Penalty with the long-term perspective of RL. Simulation results for multiple problems demonstrate that LDPTRLQ outperforms the baseline methods using the Lyapunov drift-plus-penalty method and RL, corroborating the validity of our theoretical derivations. The results also demonstrate that our proposed algorithm outperforms other benchmarks in terms of compatibility and stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。