arXiv:2507.10452cs.LG2025-07被引 1

分析连续时间LQR策略优化中的梯度主导性,揭示收敛速度与初始条件的关系。

Some remarks on gradient dominance and LQR policy optimization

  • 基于广义PL不等式研究梯度下降的收敛机制
  • 发现连续时间LQR存在全局线性/局部指数混合收敛特性
  • 适用于关注鲁棒性与误差影响的控制与强化学习研究者

优化问题(包括强化学习中的策略优化)通常依赖于梯度下降的变体。近年来,机器学习、控制与优化领域广泛使用Polyak-Łojasiewicz不等式(PLI)来证明损失函数在梯度流下以指数速率收敛至极小值(即数值分析中的局部线性收敛)。然而,对于连续时间LQR问题的策略迭代,该收敛速率在大初始条件下消失,导致全局线性与局部指数收敛并存。这与离散时间LQR的全局指数收敛形成鲜明对比。这一差距促使人们寻求多种类似PLI的推广条件。本报告将探讨此主题,并分析梯度估计误差(如对抗攻击、模拟错误、早期停止、近似数字孪生、随机计算或数据受限学习)对系统的影响,采用‘输入到状态稳定性’(ISS)进行分析。第二部分讨论反馈控制中‘线性前馈神经网络’的收敛性与PLI-like性质。相关工作由Arthur Castello B. de Oliveira、Leilei Cui、Zhong-Ping Jiang 和 Milad Siami 共同完成。

原文摘要 · Abstract (English)

Solutions of optimization problems, including policy optimization in reinforcement learning, typically rely upon some variant of gradient descent. There has been much recent work in the machine learning, control, and optimization communities applying the Polyak-Łojasiewicz Inequality (PLI) to such problems in order to establish an exponential rate of convergence (a.k.a. ``linear convergence'' in the local-iteration language of numerical analysis) of loss functions to their minima under the gradient flow. Often, as is the case of policy iteration for the continuous-time LQR problem, this rate vanishes for large initial conditions, resulting in a mixed globally linear / locally exponential behavior. This is in sharp contrast with the discrete-time LQR problem, where there is global exponential convergence. That gap between CT and DT behaviors motivates the search for various generalized PLI-like conditions, and this talk will address that topic. Moreover, these generalizations are key to understanding the transient and asymptotic effects of errors in the estimation of the gradient, errors which might arise from adversarial attacks, wrong evaluation by an oracle, early stopping of a simulation, inaccurate and very approximate digital twins, stochastic computations (algorithm ``reproducibility''), or learning by sampling from limited data. We describe an ``input to state stability'' (ISS) analysis of this issue. The second part discusses convergence and PLI-like properties of ``linear feedforward neural networks'' in feedback control. Much of the work described here was done in collaboration with Arthur Castello B. de Oliveira, Leilei Cui, Zhong-Ping Jiang, and Milad Siami.

LQR梯度下降收敛分析控制理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。