arXiv:2507.09026math.OCcs.LG2025-07

提出新参数化方法,实现LQG问题的全局收敛。

On the Gradient Domination of the LQG Problem

  • 用历史输入输出数据参数化控制器,替代传统动态控制器。
  • 证明了代价函数具有梯度主导性,支持全局收敛。
  • 适用于模型已知或未知场景,适合控制优化研究者。

本文研究通过策略梯度(PG)方法求解线性二次高斯(LQG)调节器问题。尽管PG在求解线性二次调节器(LQR)问题时已有坚实的理论保证,但其在LQG设置下的理论理解仍不充分。特别地,经典参数化下LQG问题缺乏梯度主导性,阻碍了全局收敛的证明。本文采用稳定控制器集合的替代参数化,并引入提升论证方法,将控制器参数化为前p个时间步的输入与输出历史数据,称为历史表示。该表示使我们能够建立LQG代价函数的梯度主导性和近似光滑性。我们在模型已知和模型未知设置下,证明了策略梯度LQG的全局收敛性与每步迭代稳定性。数值实验在开环不稳定系统上验证了全局收敛性,并展示了不同历史长度对收敛的影响。

原文摘要 · Abstract (English)

We consider solutions to the linear quadratic Gaussian (LQG) regulator problem via policy gradient (PG) methods. Although PG methods have demonstrated strong theoretical guarantees in solving the linear quadratic regulator (LQR) problem, despite its nonconvex landscape, their theoretical understanding in the LQG setting remains limited. Notably, the LQG problem lacks gradient dominance in the classical parameterization, i.e., with a dynamic controller, which hinders global convergence guarantees. In this work, we study PG for the LQG problem by adopting an alternative parameterization of the set of stabilizing controllers and employing a lifting argument. We refer to this parameterization as a history representation of the control input as it is parameterized by past input and output data from the previous p time-steps. This representation enables us to establish gradient dominance and approximate smoothness for the LQG cost. We prove global convergence and per-iteration stability guarantees for policy gradient LQG in model-based and model-free settings. Numerical experiments on an open-loop unstable system are provided to support the global convergence guarantees and to illustrate convergence under different history lengths of the history representation.

控制优化策略梯度梯度主导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。