arXiv:2604.17457math.OCcs.AI2026-04被引 2

揭示强化学习中策略快速最优化的几何机制

Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration

论文配图:Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration
图 1 · 摘自论文原文
  • 将价值迭代视为切换系统,分析最优动作类的有限时间收敛
  • 在特定条件下,策略识别速度超越经典收缩率,可达指数级
  • 适用于关注策略收敛而非值函数精确解的研究者

Q值迭代(Q-VI)通常通过贝尔曼算子的γ-压缩性进行分析,该方法仅粗略说明诱导贪婪策略何时达到最优。本文将折扣Q-VI视为切换系统,聚焦于实际最优解集(POSS),即其打破平局后的贪婪策略为最优的Q函数集合。主要结果表明,Q-VI在有限时间内进入围绕 \\(\mathcal X_1 = Q^* + \operatorname{span}(\mathbf 1)\\) 的不变管区,该区域包含于POSS中。对任意 \\(\varepsilon > 0\\),到 \\(\mathcal X_1\\) 的距离满足指数界 \\( (\bar\rho + \varepsilon)^k \\), 其中 \\(\bar\rho\\) 为投影切换族在垂直于 \\(\mathcal X_1\\) 方向上的联合谱半径。当 \\(\bar\rho < \gamma\\) 时,横向收敛速度超过经典收缩率。分析将快速策略识别与后续收敛至 \\(Q^*\\) 分离,后者仍可能受全1模式支配。同时给出了 \\(\bar\rho < \gamma\\) 成立或不成立的谱条件与图论条件。

原文摘要 · Abstract (English)

Q-value iteration (Q-VI) is usually analyzed through the \(γ\)-contraction of the Bellman operator. This argument proves convergence to \(Q^*\), but it gives only a coarse account of when the induced greedy policy becomes optimal. We study discounted Q-VI as a switching system and focus on the practically optimal solution set (POSS), the set of \(Q\)-functions whose tie-broken greedy policies are optimal. The main result shows that Q-VI reaches the optimal action class in finite time by entering an invariant tube around \(\mathcal X_1=Q^*+\operatorname{span}(\mathbf 1)\), which is contained in the POSS. For every \(\varepsilon>0\), the distance to \(\mathcal X_1\) satisfies an exponential bound with rate \((\barρ+\varepsilon)^k\), where \(\barρ\) is the joint spectral radius of the projected switching family restricted to directions transverse to \(\mathcal X_1\). When \(\barρ<γ\), this transverse convergence is faster than the classical contraction rate. The analysis separates fast policy identification from the subsequent convergence to \(Q^*\), which may still be governed by the all-ones mode. We also give spectral and graph-theoretic conditions under which the strict inequality \(\barρ<γ\) holds or fails.

强化学习价值迭代策略收敛几何分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。