解析投影贝尔曼方程的理论性质及两种求解算法的收敛条件。
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
- 基于严格负对角占优假设,证明投影贝尔曼方程有解。
- 揭示线性Q学习与近似值迭代的收敛性关系。
- 分析ε-贪心策略下解的特性,适用于强化学习理论研究者。
本文研究投影贝尔曼方程(PBE)的理论性质及其两种求解算法:线性Q学习和近似值迭代(AVI)。我们提出了两个保证PBE存在解的充分条件:严格负对角占优(SNRDD)假设,以及受AVI收敛性启发的另一条件。SNRDD假设同时确保线性Q学习的收敛性,并进一步探讨了其与AVI收敛性的关系。最后,当采用ε-贪心策略时,对PBE解的若干有趣性质进行了分析。
原文摘要 · Abstract (English)
In this paper, we study the theoretical properties of the projected Bellman equation (PBE) and two algorithms to solve this equation: linear Q-learning and approximate value iteration (AVI). We consider two sufficient conditions for the existence of a solution to PBE : strictly negatively row dominating diagonal (SNRDD) assumption and a condition motivated by the convergence of AVI. The SNRDD assumption also ensures the convergence of linear Q-learning, and its relationship with the convergence of AVI is examined. Lastly, several interesting observations on the solution of PBE are provided when using $ε$-greedy policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。