arXiv:2604.00179eess.SYcs.LG2026-04

提出常数步长下投影双时间尺度算法的有限时间误差分析方法。

Finite-Time Analysis of Projected Two-Time-Scale Stochastic Approximation

论文配图:Finite-Time Analysis of Projected Two-Time-Scale Stochastic Approximation
图 1 · 摘自论文原文
  • 基于投影与Polyak-Ruppert平均,建立可分解误差界
  • 误差由近似误差与统计误差组成,分别对应子空间和平均窗口
  • 理论清晰分离子空间选择与平均时长的影响,适用于强化学习

我们研究了在常数步长和Polyak-Ruppert平均下,带有投影的线性双时间尺度随机逼近的有限时间收敛性。建立了显式的均方误差上界,将其分解为两个可解释的分量:由约束子空间决定的近似误差,以及以亚线性速率衰减的统计误差。误差常数通过受限稳定性裕度和耦合可逆性条件表示,清晰地将子空间选择(近似误差)的影响与平均窗口长度(统计误差)的影响分离开来。通过在合成数据和强化学习问题上的多个数值实验,验证了理论结果的有效性。

原文摘要 · Abstract (English)

We study the finite-time convergence of projected linear two-time-scale stochastic approximation with constant step sizes and Polyak--Ruppert averaging. We establish an explicit mean-square error bound, decomposing it into two interpretable components, an approximation error determined by the constrained subspace and a statistical error decaying at a sublinear rate, with constants expressed through restricted stability margins and a coupling invertibility condition. These constants cleanly separate the effect of subspace choice (approximation errors) from the effect of the averaging horizon (statistical errors). We illustrate our theoretical results through a number of numerical experiments on both synthetic and reinforcement learning problems.

优化算法随机逼近强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。