arXiv:2411.01302cs.LGmath.OC2024-11被引 11

研究随机控制中探索性策略改进与Q-learning的收敛性,给出指数级收敛和量化误差边界。

Regret of exploratory policy improvement and $q$-learning

  • 在模型参数满足增长与光滑性条件下,证明探索性策略改进指数收敛。
  • 在函数逼近与随机近似动态假设下,给出Q-learning的量化误差与遗憾界。
  • 适用于强化学习中连续状态空间的理论分析,适合关注收敛性证明的研究者。

我们研究了由Jia和Zhou(J. Mach. Learn. Res., 24 (2023), 161)提出的用于控制扩散过程的Q-learning及相关算法的收敛性。对于探索性策略改进,在模型参数满足增长与正则性假设的前提下,建立了指数收敛性。对于Q-learning,基于对函数逼近和相关随机逼近动态的额外假设,推导出定量的误差与遗憾边界。

原文摘要 · Abstract (English)

We study the convergence of $q$-learning and related algorithms introduced by Jia and Zhou (J. Mach. Learn. Res., 24 (2023), 161) for controlled diffusion processes. For exploratory policy improvement, we establish exponential convergence under growth and regularity assumptions on the model parameters. For q-learning, we derive quantitative error and regret bounds under additional assumptions on the function approximation and the associated stochastic approximation dynamics.

强化学习收敛性Q-learning

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。