揭示Q-learning在非标准设定下收敛需更强结构条件
Optimistic Training and Convergence of Q-Learning -- Extended Version
- 证明理想基下仍可能多解,需额外约束保证收敛
- 发现盲目策略训练时算法不稳且无解或有多个解
- 澄清线性函数近似中收敛的必要条件
近期研究表明,在(ε,κ)-温和吉布斯策略下,使用线性函数近似的Q-learning是稳定的(参数估计有界),并存在投影贝尔曼方程(PBE)的解;其中κ为逆温度,ε>0用于增强探索。但解的唯一性及在标准表格或线性马尔可夫决策过程之外的收敛条件仍未解决。本文拓展了这些结果,分析了其他变体:一维示例显示,在盲目策略训练下,可能不存在PBE解或存在多个解,此时算法不稳定。主要贡献在于揭示收敛需要远更严格的结构条件。一个例子表明,尽管基是理想的(真实Q函数属于基的张成空间),但在贪婪策略下,PBE仍有两解,因此对所有足够小的ε>0和κ≥1的(ε,κ)-温和吉布斯策略也存在两个解。
原文摘要 · Abstract (English)
In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the $(\varepsilon,κ)$-tamed Gibbs policy; $κ$ is inverse temperature, and $\varepsilon>0$ is introduced for additional exploration. Under these assumptions it also follows that there is a solution to the projected Bellman equation (PBE). Left open is uniqueness of the solution, and criteria for convergence outside of the standard tabular or linear MDP settings. The present work extends these results to other variants of Q-learning, and clarifies prior work: a one dimensional example shows that under an oblivious policy for training there may be no solution to the PBE, or multiple solutions, and in each case the algorithm is not stable under oblivious training. The main contribution is that far more structure is required for convergence. An example is presented for which the basis is ideal, in the sense that the true Q-function is in the span of the basis. However, there are two solutions to the PBE under the greedy policy, and hence also for the $(\varepsilon,κ)$-tamed Gibbs policy for all sufficiently small $\varepsilon>0$ and $κ\ge 1$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。