arXiv:2502.14208cs.LGmath.OC2025-02被引 11

提出非渐近稳定理论,统一分析迭代算法在确定与随机情形下的收敛性。

A Non-Asymptotic Theory of Seminorm Lyapunov Stability: From Deterministic to Stochastic Iterative Algorithms

  • 基于半范数压缩算子,建立迭代算法的几何收敛机制。
  • 在随机情形下,无需赫尔维茨条件即实现线性马尔可夫型算法的有限样本分析。
  • 适用于平均奖励强化学习中的TD(λ)和Q-learning,提供统一理论框架。

本文研究半范数压缩算子的不动点方程求解问题,建立确定与随机迭代算法的非渐近行为基础理论。在确定情形下,证明了半范数压缩算子的不动点定理,表明迭代序列以几何速率收敛至半范数的核空间。在随机情形下,分析了在半范数压缩算子与马尔可夫噪声下的随机逼近(SA)算法,针对多种步长选择提供了有限样本分析。以线性方程组为基准,其不动点迭代的收敛性与线性动力系统的稳定性密切相关;在此特例中,我们的结果完整刻画了关于半范数的系统稳定性,将其与正半定矩阵的李雅普诺夫方程解相联系。在随机情形下,我们建立了无赫尔维茨假设的线性马尔可夫型SA的有限样本分析。理论结果为平均奖励设置下多种强化学习算法的有限样本界推导提供了统一框架,包括策略评估的TD(λ)(作为泊松方程求解的特例)和控制的Q-learning。

原文摘要 · Abstract (English)

We study the problem of solving fixed-point equations for seminorm-contractive operators and establish foundational results on the non-asymptotic behavior of iterative algorithms in both deterministic and stochastic settings. Specifically, in the deterministic setting, we prove a fixed-point theorem for seminorm-contractive operators, showing that iterates converge geometrically to the kernel of the seminorm. In the stochastic setting, we analyze the corresponding stochastic approximation (SA) algorithm under seminorm-contractive operators and Markovian noise, providing a finite-sample analysis for various stepsize choices. A benchmark for equation solving is linear systems of equations, where the convergence behavior of fixed-point iteration is closely tied to the stability of linear dynamical systems. In this special case, our results provide a complete characterization of system stability with respect to a seminorm, linking it to the solution of a Lyapunov equation in terms of positive semi-definite matrices. In the stochastic setting, we establish a finite-sample analysis for linear Markovian SA without requiring the Hurwitzness assumption. Our theoretical results offer a unified framework for deriving finite-sample bounds for various reinforcement learning algorithms in the average reward setting, including TD($λ$) for policy evaluation (which is a special case of solving a Poisson equation) and Q-learning for control.

优化理论强化学习收敛分析随机逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。