arXiv:2508.21314cs.LG2025-08中稿 · CDC 2025

提出一种新框架,证明带状态记忆和正则化的强化学习能收敛到稳定解。

Convergence of regularized agent-state-based Q-learning in POMDPs

  • 用非信念状态的递归网络状态更新Q表,结合策略正则化提升稳定性。
  • 在温和条件下,算法收敛至由行为策略决定的正则化马尔可夫决策过程的固定点。
  • 适用于研究强化学习收敛性或设计更稳定的智能体策略的学者。

本文提出一个框架,用于理解实践中常用Q-learning强化学习算法的收敛性。这类算法的两个显著特征是:(i) Q表通过非信念状态或信息状态的智能体状态(如循环神经网络状态)进行递归更新;(ii) 常使用策略正则化以促进探索并稳定学习过程。我们研究此类算法的最简形式,称为正则化代理状态基础Q-learning(RASQL),并证明其在温和技术条件下收敛至一个恰当定义的正则化马尔可夫决策过程(MDP)的固定点,该固定点依赖于行为策略诱导的平稳分布。我们还表明,该分析同样适用于学习周期策略的RASQL变体。数值实验验证了经验收敛行为与所提出的理论极限一致。

原文摘要 · Abstract (English)

In this paper, we present a framework to understand the convergence of commonly used Q-learning reinforcement learning algorithms in practice. Two salient features of such algorithms are: (i)~the Q-table is recursively updated using an agent state (such as the state of a recurrent neural network) which is not a belief state or an information state and (ii)~policy regularization is often used to encourage exploration and stabilize the learning algorithm. We investigate the simplest form of such Q-learning algorithms which we call regularized agent-state-based Q-learning (RASQL) and show that it converges under mild technical conditions to the fixed point of an appropriately defined regularized MDP, which depends on the stationary distribution induced by the behavioral policy. We also show that a similar analysis continues to work for a variant of RASQL that learns periodic policies. We present numerical examples to illustrate that the empirical convergence behavior matches with the proposed theoretical limit.

强化学习收敛性POMDPQ-learning

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。