arXiv:2510.17391cs.LG2025-10NeurIPS被引 1

首个针对平均奖励强化学习的有限时间分析,突破了传统假设限制。

Finite-Time Bounds for Average-Reward Fitted Q-Iteration

  • 提出锚定拟合Q迭代,通过锚点机制实现稳定学习
  • 在弱连通MDP下证明了样本复杂度的有限时间上界
  • 适用于单轨迹数据,适合实际应用中的离线强化学习

尽管已有大量研究关注带函数逼近的折扣回报离线强化学习的样本复杂度,但针对平均奖励设置的研究仍较少,且现有方法依赖于如遍历性或线性MDP等强假设。本文首次为弱连通MDP下的平均奖励离线强化学习建立了样本复杂度结果。为此,我们提出了锚定拟合Q迭代(Anchored Fitted Q-Iteration),结合标准拟合Q迭代与锚点机制。我们证明,该锚点可视为一种权重衰减形式,对实现平均奖励设置下的有限时间分析至关重要。此外,我们的有限时间分析还扩展至数据由单条轨迹生成而非独立同分布转移的情形,同样依赖锚点机制。

原文摘要 · Abstract (English)

Although there is an extensive body of work characterizing the sample complexity of discounted-return offline RL with function approximations, prior work on the average-reward setting has received significantly less attention, and existing approaches rely on restrictive assumptions, such as ergodicity or linearity of the MDP. In this work, we establish the first sample complexity results for average-reward offline RL with function approximation for weakly communicating MDPs, a much milder assumption. To this end, we introduce Anchored Fitted Q-Iteration, which combines the standard Fitted Q-Iteration with an anchor mechanism. We show that the anchor, which can be interpreted as a form of weight decay, is crucial for enabling finite-time analysis in the average-reward setting. We also extend our finite-time analysis to the setup where the dataset is generated from a single-trajectory rather than IID transitions, again leveraging the anchor mechanism.

强化学习离线学习样本复杂度平均奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。