用切换线性系统分析线性Q-learning,给出收敛性更优的理论保证。
A Switching System Theory of Q-Learning with Linear Function Approximation

- 将线性Q-learning误差建模为随机切换线性系统。
- 基于联合谱半径得到有限时间误差上界,刻画最坏情况衰减速率。
- 提出新收敛判据,比传统一步范数方法更宽松有效。
Q-learning是强化学习中的基础算法。本文从切换线性系统(SLS)视角构建线性Q-learning的新分析框架,其中线性Q-learning指采用线性函数逼近的Q-learning。我们推导出线性Q-learning误差的随机SLS表示,并通过相关SLS族的联合谱半径(JSR)获得线性Q-learning的有限时间误差分析;JSR即对应SLS在最坏情况下的指数衰减速率。基于JSR的速率精确关联于该SLS表示的内在最坏情况指数速率。此外,我们给出了一个基于JSR的线性Q-learning收敛性证书,其保守性低于传统的单步范数界限。
原文摘要 · Abstract (English)
Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing linear Q-learning from a switching linear system (SLS) viewpoint, where linear Q-learning denotes Q-learning with linear function approximation. We derive a stochastic SLS representation of the linear Q-learning error and obtain a finite-time error analysis for linear Q-learning through the joint spectral radius (JSR) of the associated SLS family; the JSR is the exact worst-case exponential rate of the corresponding SLSs. The JSR-based rate is tied to the intrinsic worst-case exponential rate of the SLS representation. Moreover, we provide a JSR-based certificate for convergence of linear Q-learning, which can be less conservative than one-step norm bounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。