arXiv:2505.21391cs.LGcs.AI2025-05NeurIPS被引 6

首次给出任意特征下线性TD(λ)的收敛速率分析

Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary Features

  • 在任意特征下分析线性TD(λ)的收敛性,无需修改算法或额外假设
  • 首次建立$ L^2 $收敛速率,适用于折扣与平均奖励两种设置
  • 提出新随机逼近理论,收敛至解集而非唯一解,适合特征冗余场景

线性TD(λ)是策略评估中最基础的强化学习算法之一。以往的收敛速率分析通常依赖于特征线性无关的假设,但这一假设在许多实际场景中不成立。本文首次在任意特征条件下建立了线性TD(λ)的$ L^2 $收敛速率,无需算法修改或额外假设。结果同时适用于折扣回报和平均奖励设定。针对任意特征可能导致解不唯一的问题,我们提出了一个新的随机逼近结论,其收敛速率指向解集而非单一点。

原文摘要 · Abstract (English)

Linear TD($λ$) is one of the most fundamental reinforcement learning algorithms for policy evaluation. Previously, convergence rates are typically established under the assumption of linearly independent features, which does not hold in many practical scenarios. This paper instead establishes the first $L^2$ convergence rates for linear TD($λ$) operating under arbitrary features, without making any algorithmic modification or additional assumptions. Our results apply to both the discounted and average-reward settings. To address the potential non-uniqueness of solutions resulting from arbitrary features, we develop a novel stochastic approximation result featuring convergence rates to the solution set instead of a single point.

强化学习TD学习收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。