arXiv:2510.23914cs.LG2025-10被引 1

统一分析价值迭代收敛性,发现两种奖励情形下均几何收敛且更快。

Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases

  • 基于几何视角统一分析折扣与平均奖励情形下的价值迭代
  • 在唯一最优策略且无环条件下,收敛速度比已有理论更快
  • 适合研究强化学习理论或算法收敛性的研究人员

尽管价值迭代(VI)是强化学习中最基础的算法之一,其理论收敛保证仍与实际表现存在明显差距。在折扣奖励情形下,经典理论保证以速率γ进行几何收敛;而在平均奖励情形下,近期研究认为仅能实现次线性收敛。然而实践中,VI通常收敛得快得多。本文通过统一的几何分析表明,在存在唯一且无环最优策略的假设下:(i) VI在折扣与平均奖励设定中均为几何收敛;(ii) 收敛速率优于先前理论预测。

原文摘要 · Abstract (English)

While Value Iteration (VI) is one of the most fundamental algorithms in Reinforcement Learning, its theoretical convergence guarantees still exhibit a persistent mismatch with empirical behavior. In the discounted-reward case, classical theory guarantees geometric convergence with rate $γ$, while in the average-reward case recent work suggests that only sublinear convergence can be expected. In practice, however, VI is often observed to converge significantly faster. In this work, we show through a unified geometry-based analysis that, under an assumption of a unique and unichain optimal policy, (i) convergence is geometric in both the discounted- and average-reward settings and (ii) the convergence rate is faster than previous analyses suggest.

强化学习收敛性价值迭代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。