优化量化器提升无限状态马尔可夫决策过程的近似精度
Quantizer Design for Finite Model Approximations, Model Learning, and Quantized Q-Learning for MDPs with Unbounded Spaces
- 通过优化量化器设计,降低无限状态MDP的有限模型近似误差
- 在李雅普诺夫增长条件下,误差随分箱数增加趋于零
- 适用于需量化状态的Q-learning与模型学习场景
针对具有无界状态空间的马尔可夫决策过程(MDPs),本文在[Kara等,JMLR'23]的上界基础上,通过优化用于有限模型近似的量化器,给出了更精细的上界。研究还探讨了量化器设计在量化Q-learning与经验模型学习中的影响,以及在将量化状态视为真实状态时所获得策略的性能表现。文章强调了规划与学习(包括Q-learning或经验模型学习)之间的本质区别:规划中可独立设计近似MDP,而学习中近似MDP必须由探索策略下的马尔可夫链不变测度决定,导致量化器设计性能存在显著差异,尽管两种情形下均可实现渐近近似最优。特别地,在李雅普诺夫增长条件下,获得了显式的上界,其随分箱数趋于无穷而收敛至零。
原文摘要 · Abstract (English)
In this paper, for Markov decision processes (MDPs) with unbounded state spaces we present refined upper bounds presented in [Kara et. al. JMLR'23] on finite model approximation errors via optimizing the quantizers used for finite model approximations. We also consider implications on quantizer design for quantized Q-learning and empirical model learning, and the performance of policies obtained via Q-learning where the quantized state is treated as the state itself. We highlight the distinctions between planning, where approximating MDPs can be independently designed, and learning (either via Q-learning or empirical model learning), where approximating MDPs are restricted to be defined by invariant measures of Markov chains under exploration policies, leading to significant subtleties on quantizer design performance, even though asymptotic near optimality can be established under both setups. In particular, under Lyapunov growth conditions, we obtain explicit upper bounds which decay to zero as the number of bins approaches infinity
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。