arXiv:2509.09863eess.SYcs.LG2025-09被引 2

提出离策略学习李雅普诺夫函数,提升强化学习算法的稳定性和数据效率。

Off Policy Lyapunov Stability in Reinforcement Learning

  • 离策略学习李雅普诺夫函数,突破传统方法依赖在线采样的瓶颈。
  • 在倒立摆和四旋翼系统上验证,显著提升SAC和PPO的性能与稳定性。
  • 适合关注强化学习稳定性与高效训练的研究者和工程应用开发者。

传统强化学习缺乏稳定性保障。近期算法通过自学习李雅普诺夫函数来确保学习过程稳定,但现有方法因采用在线学习方式导致样本效率低下。本文提出一种离策略学习李雅普诺夫函数的方法,并将其集成至Soft Actor Critic(SAC)和Proximal Policy Optimization(PPO)算法中,为二者提供高效的数据利用型稳定性证明。通过倒立摆与四旋翼系统的仿真验证,引入该方法后,SAC与PPO在稳定性与收敛性方面均有显著提升。

原文摘要 · Abstract (English)

Traditional reinforcement learning lacks the ability to provide stability guarantees. More recent algorithms learn Lyapunov functions alongside the control policies to ensure stable learning. However, the current self-learned Lyapunov functions are sample inefficient due to their on-policy nature. This paper introduces a method for learning Lyapunov functions off-policy and incorporates the proposed off-policy Lyapunov function into the Soft Actor Critic and Proximal Policy Optimization algorithms to provide them with a data efficient stability certificate. Simulations of an inverted pendulum and a quadrotor illustrate the improved performance of the two algorithms when endowed with the proposed off-policy Lyapunov function.

强化学习稳定性离策略李雅普诺夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。