arXiv:2409.08382eess.SYcs.LG2024-09被引 2

通过学习局部线性模型实现未知非线性系统的稳定控制。

Stochastic Reinforcement Learning with Stability Guarantees for Control of Unknown Nonlinear Systems

  • 用神经网络学习系统局部线性表示,将增益矩阵直接融入策略。
  • 在高维系统上优于SAC和PPO,实现渐近稳定性。
  • 理论分析支持算法可行性,适合复杂控制系统设计者。

为未知非线性系统设计稳定控制器是极具挑战性的任务,尤其在高维情形下。传统强化学习方法常使系统趋近平衡点但无法真正稳定,导致持续振荡。本文提出一种强化学习算法,通过学习系统动态的局部线性表示,并将学习到的增益矩阵直接集成到神经策略中。在多个高维动力系统上的仿真表明,该算法性能优于SAC与PPO等主流算法,成功实现系统稳定。进一步地,本文提供了针对确定性和随机强化学习设置的可行性分析,以及所提算法的收敛性证明。结果验证了学习到的控制策略确实能为非线性系统提供渐近稳定性。

原文摘要 · Abstract (English)

Designing a stabilizing controller for nonlinear systems is a challenging task, especially for high-dimensional problems with unknown dynamics. Traditional reinforcement learning algorithms applied to stabilization tasks tend to drive the system close to the equilibrium point. However, these approaches often fall short of achieving true stabilization and result in persistent oscillations around the equilibrium point. In this work, we propose a reinforcement learning algorithm that stabilizes the system by learning a local linear representation ofthe dynamics. The main component of the algorithm is integrating the learned gain matrix directly into the neural policy. We demonstrate the effectiveness of our algorithm on several challenging high-dimensional dynamical systems. In these simulations, our algorithm outperforms popular reinforcement learning algorithms, such as soft actor-critic (SAC) and proximal policy optimization (PPO), and successfully stabilizes the system. To support the numerical results, we provide a theoretical analysis of the feasibility of the learned algorithm for both deterministic and stochastic reinforcement learning settings, along with a convergence analysis of the proposed learning algorithm. Furthermore, we verify that the learned control policies indeed provide asymptotic stability for the nonlinear systems.

强化学习非线性控制稳定性神经控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。