通过几何分析揭示强化学习中值迭代的收敛更快原因
Switching-Geometry Analysis of Deflated Q-Value Iteration

- 用切换系统几何方法分析去偏值迭代算法
- 发现去偏后收敛速率可低于原折扣因子γ
- 适合研究强化学习收敛性与算法优化的读者
本文构建了联合谱半径(JSR)框架,用于分析折扣马尔可夫决策过程控制中的秩一去偏值迭代(deflated Q-VI)。聚焦于全1残差修正,通过切换系统几何视角进行解释,并首次对政策优化问题中的去偏Q-VI给出了基于JSR的收敛性分析。分析表明,标准Q-VI切换系统模型的JSR恰好等于折扣因子γ∈(0,1),因为所有允许的子系统均共享全1向量作为不变方向。通过移除该方向的商空间投影,得到一个新投影切换系统模型,其JSR决定相关误差动态,可能严格小于γ。因此,去偏Q-VI可获得比原γ界更精细的收敛速率刻画。最后证明,该修正等价于标准Q-VI的标量重中心化,故投影轨迹与贪婪策略序列与同初始点的标准Q-VI一致。去偏的优势不在于改变决策问题,而在于去除冗余全1分量后对收敛几何的更精确描述。
原文摘要 · Abstract (English)
This paper develops a joint spectral radius (JSR) framework for analyzing rank-one deflated Q-value iteration (Q-VI) in discounted Markov decision process control. Focusing on an all-ones residual correction, we interpret the resulting algorithm through the geometry of switching systems and, to the best of our knowledge, give the first JSR-based convergence analysis of deflated Q-VI for policy optimization problems. Our analysis reveals that the standard Q-VI switching system model has JSR exactly the discount factor $γ\in (0,1)$, since all admissible subsystems share the all-ones vector as an invariant direction. By passing to the quotient space that removes this direction, we obtain a projected switching system model whose JSR governs the relevant error dynamics and may be strictly smaller than $γ$. Therefore, the deflated Q-VI admits a potentially sharper convergence-rate characterization than the ambient-space $γ$-bound. Finally, we prove that the correction is equivalent to a scalar recentering of standard Q-VI. Hence, the projected trajectory, and therefore the greedy-policy sequence, is unchanged relative to standard Q-VI initialized from the same point. The benefit of deflation is not a change in the induced decision-making problem, but a more precise JSR-based description of the convergence geometry after the redundant all-ones component is removed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。