为连续控制的后验采样强化学习提供首个无界状态空间下的次线性后悔界。
Posterior Sampling Reinforcement Learning with Gaussian Processes for Continuous Control: Sublinear Regret Bounds for Unbounded State Spaces
- 利用高斯过程后验采样,在弱光滑条件下设计算法并分析其性能。
- 证明了后悔率可达 $ ilde{O}(H oot{H}{γ_T T})$,实现次线性增长。
- 适用于复杂动态系统建模,为理论分析提供新工具。
我们分析了高斯过程后验采样强化学习(GP-PSRL)的贝叶斯后悔。后验采样是一种在不确定性下决策的启发式方法,已被用于解决多种连续控制问题。然而,关于GP-PSRL的理论研究仍有限:已知的后悔界或存在次优增长速率、依赖强光滑性假设,或未能正确处理状态空间无界的事实。通过递归应用Borell-Tsirelson-Ibragimov-Sudakov不等式,我们证明,以高概率,算法实际访问的状态被限制在一个近似恒定半径的球内。随后,采用链式方法在弱光滑条件下控制GP-PSRL的后悔。主要结果是贝叶斯后悔界为 $ ilde{O}(H oot{H}{γ_T T})$,其中 $H$ 为时间步长,$T$ 为总时间步数,$γ_T$ 为期望信息增益。该结果解决了先前理论工作的局限性,并为复杂场景下PSRL的分析提供了理论基础与工具。
原文摘要 · Abstract (English)
We analyze the Bayesian regret of the Gaussian process posterior sampling reinforcement learning (GP-PSRL) algorithm. Posterior sampling is a heuristic for decision-making under uncertainty that has been used to develop successful algorithms for a variety of continuous control problems. However, theoretical work on GP-PSRL is limited. All known regret bounds either have a sub-optimal growth rate, require strong smoothness assumptions, or fail to properly account for the fact that the set of possible system states is unbounded. Through a recursive application of the Borell-Tsirelson-Ibragimov-Sudakov inequality, we show that, with high probability, the states actually visited by the algorithm are contained within a ball of near-constant radius. We then use the chaining method to control the regret suffered by GP-PSRL under weak smoothness conditions. Our main result is a Bayesian regret bound of the order $\widetilde{\mathcal{O}}(H\sqrt{γ_TT})$, where $H$ is the horizon, $T$ is the number of time steps and $γ_T$ is the expected information gain. With this result, we resolve the limitations with prior theoretical work on PSRL, and provide the theoretical foundation and tools for analyzing PSRL in complex settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。