arXiv:2410.13067eess.SYcs.LG2024-10被引 5

常步长两尺度随机逼近实现稳定收敛,精度提升显著。

Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way

  • 用马尔可夫过程分析常步长迭代,揭示收敛机制
  • 偏置与步长线性相关,方差仅随自身步长变化
  • 无需强假设,适合强化学习等复杂场景

以往对两尺度随机逼近的研究多聚焦于递减步长下的均方误差界。本文从马尔可夫过程视角出发,研究常步长方案,证明两个时间尺度的迭代序列在Wasserstein度量下收敛至唯一联合平稳分布。我们推导出明确的几何与非渐近收敛速率,以及常步长在马尔可夫噪声下的方差与偏置。具体而言,当使用两个常步长α<β时,偏置以Θ(α)+Θ(β)阶增长(含高阶项),而慢速迭代(对应α)的方差为O(α),快速迭代(对应β)的方差为O(β)。与以往工作不同,本结果无需β²≪α或维度依赖等额外假设。精细刻画使尾部平均与外推技术可有效降低方差与偏置,将均方误差界优化至O(β⁴ + 1/t),适用于两个迭代序列。

原文摘要 · Abstract (English)

Previous studies on two-timescale stochastic approximation (SA) mainly focused on bounding mean-squared errors under diminishing stepsize schemes. In this work, we investigate {\it constant} stpesize schemes through the lens of Markov processes, proving that the iterates of both timescales converge to a unique joint stationary distribution in Wasserstein metric. We derive explicit geometric and non-asymptotic convergence rates, as well as the variance and bias introduced by constant stepsizes in the presence of Markovian noise. Specifically, with two constant stepsizes $α< β$, we show that the biases scale linearly with both stepsizes as $Θ(α)+Θ(β)$ up to higher-order terms, while the variance of the slower iterate (resp., faster iterate) scales only with its own stepsize as $O(α)$ (resp., $O(β)$). Unlike previous work, our results require no additional assumptions such as $β^2 \ll α$ nor extra dependence on dimensions. These fine-grained characterizations allow tail-averaging and extrapolation techniques to reduce variance and bias, improving mean-squared error bound to $O(β^4 + \frac{1}{t})$ for both iterates.

随机逼近常步长收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。