新算法让强化学习在复杂环境下更高效地估算价值函数。
Planning in entropy-regularized Markov decision processes and games
- 利用熵正则化提升贝尔曼算子平滑性,改进规划效率。
- 样本复杂度达 O~(1/epsilon^4),优于传统方法无保证的多项式复杂度。
- 适合需要稳定收敛的强化学习与博弈场景研究者使用。
我们提出 SmoothCruiser,一种在熵正则化马尔可夫决策过程和双人博弈中,基于环境生成模型估计价值函数的新规划算法。该算法利用正则化带来的贝尔曼算子平滑性,实现了对精度 epsilon 的问题无关样本复杂度为 O~(1/epsilon^4)。而在非正则化设置下,目前尚无已知算法能在最坏情况下保证多项式样本复杂度。
原文摘要 · Abstract (English)
We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the environment. SmoothCruiser makes use of the smoothness of the Bellman operator promoted by the regularization to achieve problem-independent sample complexity of order O~(1/epsilon^4) for a desired accuracy epsilon, whereas for non-regularized settings there are no known algorithms with guaranteed polynomial sample complexity in the worst case.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。