arXiv:2601.19284eess.SYcs.LG2026-01

无需系统模型,用策略梯度法实现输出反馈下的系统稳定控制。

Model-Free Output Feedback Stabilization via Policy Gradient Methods

  • 基于系统轨迹的零阶策略梯度更新,实现无模型学习。
  • 算法收敛至稳定点,可求得离散线性系统的稳定输出反馈策略。
  • 适用于状态不完全可观、无模型信息的控制系统设计者。

稳定动态系统是控制系统中的基础问题,尤其在系统模型未知时更具挑战性。尽管强化学习中的策略梯度(PG)方法因实现简单且可无模型求解,已成功应用于未知线性动态系统,但现有工作多假设全状态反馈。本文首次将策略梯度方法拓展至部分可观测的输出反馈场景,聚焦于系统稳定性的基本问题。我们提出一种算法框架,通过基于系统轨迹的零阶策略梯度更新,虽无全局收敛保证,仍可收敛至平稳点,并获得离散时间线性动态系统的稳定输出反馈策略。同时,我们明确刻画了算法的样本复杂度,并通过数值实验验证了其有效性。

原文摘要 · Abstract (English)

Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning (RL) algorithms that have been successfully applied to solve problems pertaining to unknown linear dynamical systems, the policy gradient (PG) method stands out due to its ease of implementation and can solve the problem in a model-free manner. However, most of the existing works on PG methods for unknown linear dynamical systems assume full-state feedback. In this paper, we take a step towards model-free learning for partially observable linear dynamical systems with output feedback and focus on the fundamental stabilization problem of the system. We propose an algorithmic framework that stretches the boundary of PG methods to the problem without global convergence guarantees. We show that by leveraging zeroth-order PG update based on system trajectories and its convergence to stationary points, the proposed algorithms return a stabilizing output feedback policy for discrete-time linear dynamical systems. We also explicitly characterize the sample complexity of our algorithm and verify the effectiveness of the algorithm using numerical examples.

控制理论策略梯度无模型学习输出反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。