arXiv:2511.12688stat.MLcs.LG2025-11

提出高效分布强化学习算法,可在线学习回报分布而不受支持集大小影响。

Accelerated Distributional Temporal Difference Learning with Linear Function Approximation

  • 基于线性函数逼近与方差缩减技术,设计新算法。
  • 样本复杂度不随支持集大小K增长,理论更优。
  • 适合关注分布式强化学习效率的科研人员。

本文研究了带线性函数逼近的分布时序差分(Distributional TD)学习的有限样本统计速率。分布TD学习旨在估计给定策略下折扣马尔可夫决策过程的回报分布。以往工作主要聚焦于无函数逼近的表格情形。本文首先分析线性-分类贝尔曼方程,进而引入方差缩减技术,提出新算法,在$K$较大时实现了与支持集大小$K$无关的紧致样本复杂度界。理论结果表明:使用线性函数逼近的分布TD学习,从流数据中学习回报分布的难度并不高于学习其期望值。本工作为分布强化学习的统计效率提供了新见解。

原文摘要 · Abstract (English)

In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The purpose of distributional TD learning is to estimate the return distribution of a discounted Markov decision process for a given policy. Previous works on statistical analysis of distributional TD learning focus mainly on the tabular case. We first consider the linear function approximation setting and conduct a fine-grained analysis of the linear-categorical Bellman equation. Building on this analysis, we further incorporate variance reduction techniques in our new algorithms to establish tight sample complexity bounds independent of the support size $K$ when $K$ is large. Our theoretical results imply that, when employing distributional TD learning with linear function approximation, learning the full distribution of the return function from streaming data is no more difficult than learning its expectation. This work provide new insights into the statistical efficiency of distributional reinforcement learning algorithms.

强化学习分布学习线性逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。