arXiv:2608.27313stat.MLcs.LG2026-08

首次为分布强化学习的分位数TD提供有限样本保证,解析了收敛稳定性机制。

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

  • 基于分布贝尔曼算子的收缩性与奖励分布单调性,将初始值拉入局部邻域。
  • 在邻域内线性化后,证明最后迭代波动为 $\widetilde O(T^{-a/2}/\sqrt{1-γ})$,不依赖分位点数。
  • 适用于关注分布强化学习理论分析的科研人员,尤其关注收敛性与样本复杂度分离。

我们为表格型分布强化学习中的同步分位数时序差分学习(QTD)建立了全局有限样本保证。证明分离了两种稳定性机制:基于奖励累积分布函数的序单调性和分布贝尔曼算子的 $W_\infty$ 收缩性,可将任意初始化的迭代点拉入局部邻域;在该邻域内,对 QTD 的均场进行线性化,其雅可比矩阵为非奇异 $M$-矩阵,关联的正半群支持方差敏感的鞅分析。对于步长 $α_t = c(t+1)^{-a}$ 且 $a \in (1/2,1)$,主导的最后迭代波动量级为 $\widetilde O(T^{-a/2}/\sqrt{1-γ})$,且不具有分位点数量的多项式依赖。确定性暂态和所需预热阶段仍可能依赖最小贝尔曼目标密度,最坏情况下为 $m^{-1}$。因此,结果清晰区分了局部随机波动与全局样本复杂度。

原文摘要 · Abstract (English)

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes $α_t=c(t+1)^{-a}$ with $a\in(1/2,1)$, the leading last-iterate fluctuation is of order $\widetilde O\bigl(T^{-a/2}/\sqrt{1-γ}\bigr)$ and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order $m^{-1}$ in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.

强化学习分布学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。