arXiv:2604.19569cs.LGcs.AI2026-04被引 9

用切换线性系统分析Q-learning收敛速度,给出精确的误差衰减速率。

Switching Theory for Q-Learning

论文配图:Switching Theory for Q-Learning
图 1 · 摘自论文原文
  • 从切换线性系统视角建模Q-learning误差演化过程。
  • 首次用联合谱半径精确刻画常步长Q-learning的收敛速率。
  • 结果更紧致,适合研究算法收敛性与理论分析者阅读。

Q-learning是强化学习中的基础算法。本文从切换线性系统(SLS)视角,为常步长表格式Q-learning提出新分析框架。具体地,推导了Q-learning误差的随机SLS表示,并通过对应SLS模型的联合谱半径(JSR)进行有限时间误差分析,其中JSR即该SLS的精确最坏情况指数衰减速率。据我们所知,这是首个将标准Q-learning的主导指数收敛率以JSR形式表达的分析。该速率关联于SLS表示的内在最坏情况指数速率,当传统行和上界过于保守时可更紧致。进一步证明Q-learning的JSR等于所有确定性策略模式中最大谱半径,并给出可任意精度计算的线性规划刻画。

原文摘要 · Abstract (English)

Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing constant step-size tabular Q-learning from a switching linear system (SLS) viewpoint. In particular, we derive a stochastic SLS representation of the Q-learning error, and a finite-time error analysis through the joint spectral radius (JSR) of the corresponding SLS model, where the JSR is the exact worst-case exponential rate of the associated SLS. To the best of our knowledge, this is the first convergence rate analysis of standard Q-learning whose leading exponential rate is expressed through the JSR. The resulting rate is tied to the intrinsic worst-case exponential rate of the direct SLS representation and can be sharper than row-sum upper bounds when those bounds are conservative. We further prove that the JSR of Q-learning equals the largest spectral radius among the deterministic-policy modes and give an exact linear programming characterization that can be evaluated to any prescribed accuracy.

强化学习收敛分析线性系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。