arXiv:2605.06866cs.LGmath.OC2026-05被引 2

提出异步分布时序差分学习的有限迭代理论,统一分析多种场景。

A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning

  • 基于压缩映射与受限域分析,建立异步更新的收敛机制。
  • 在独立同分布与马尔可夫轨迹下,给出折扣情形的有限步误差界。
  • 适用于单标量、多变量及固定时域场景,适合强化学习研究者。

本文研究分类分布时序差分方法所用精确异步递推的有限迭代行为。分析涵盖标量分类TD(Cramér几何)与多变量带符号分类TD(最大均值差异几何)。现有状态逐点等距嵌入将两类方法转化为单状态随机逼近递推,其在块上确界范数下收缩,但分类算子仅在不变表示域上具有收缩性。本文建立所需的受限域理论,并在独立同分布采样和马尔可夫轨迹下获得折扣误差界。通过泊松方程分解处理轨迹依赖性,无需显式混合时间窗。对于无折扣固定时域策略评估,还在周期采样下建立了类似的有限迭代保证。这些结果共同提供了异步分类分布TD在标量、多变量、折扣及固定时域设置下的统一非渐近分析。

原文摘要 · Abstract (English)

We study finite-iteration behavior of the exact asynchronous recursions used by categorical distributional temporal-difference methods. The analysis covers scalar categorical TD in the Cramér geometry and multivariate signed-categorical TD in the maximum mean discrepancy geometry. Existing statewise isometric embeddings turn both methods into single-state stochastic-approximation recursions that contract in a block-supremum norm, but the categorical operators are contractive only on invariant representation domains. We establish the required restricted-domain theory and obtain discounted bounds under i.i.d. sampling and under a Markovian trajectory. A Poisson-equation decomposition handles trajectory dependence without an explicit mixing-time window. For undiscounted fixed-horizon policy evaluation, we establish analogous finite-iteration guarantees for horizon-stacked categorical methods under episodic sampling. Together, these results provide a unified non-asymptotic analysis of asynchronous categorical distributional TD across scalar, multivariate, discounted, and fixed-horizon settings.

强化学习分布学习收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。