提出更紧的代价悲观与奖励乐观策略,显著提升安全强化学习的性能上限。
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
- 基于贝尔曼型方差律设计新型函数估计器,实现更紧的代价悲观与奖励乐观
- 在不违反约束前提下,达到新上界 $\widetilde{\mathcal{O}}((\bar C - \bar C_b)^{-1}H^{2.5} S\sqrt{AK})$
- 当 $\bar C - \bar C_b=Ω(H)$ 时,逼近最优下界,适合高安全要求场景
本文研究了在未知转移核和随机奖励、代价函数下的安全强化学习问题,形式化为有限时域、表格型带约束马尔可夫决策过程。提出一种基于新估计器的模型化算法,实现更紧的代价悲观与奖励乐观。该算法在每轮均保证不违反约束,达到 $ ilde{\mathcal{O}}((\bar C - \bar C_b)^{-1}H^{2.5} S\sqrt{AK})$ 的遗憾上界,其中 $\bar C$ 为单轮成本预算,$\bar C_b$ 为安全基线策略的期望成本,$H$ 为时域长度,$S$、$A$、$K$ 分别为状态数、动作数和总轮数。该上界优于现有最佳结果;当 $\bar C - \bar C_b = Ω(H)$ 时,几乎匹配 $Ω(H^{1.5}\sqrt{SAK})$ 的遗憾下界。通过贝尔曼型全方差定律推导出代价与奖励函数估计器,获得价值函数估计期望方差和的紧界,从而降低对时域 $H$ 的依赖。实验验证了所提框架的计算有效性。
原文摘要 · Abstract (English)
This paper studies the safe reinforcement learning problem formulated as an episodic finite-horizon tabular constrained Markov decision process with an unknown transition kernel and stochastic reward and cost functions. We propose a model-based algorithm based on novel cost and reward function estimators that provide tighter cost pessimism and reward optimism. While guaranteeing no constraint violation in every episode, our algorithm achieves a regret upper bound of $\widetilde{\mathcal{O}}((\bar C - \bar C_b)^{-1}H^{2.5} S\sqrt{AK})$ where $\bar C$ is the cost budget for an episode, $\bar C_b$ is the expected cost under a safe baseline policy over an episode, $H$ is the horizon, and $S$, $A$ and $K$ are the number of states, actions, and episodes, respectively. This improves upon the best-known regret upper bound, and when $\bar C- \bar C_b=Ω(H)$, it nearly matches the regret lower bound of $Ω(H^{1.5}\sqrt{SAK})$. We deduce our cost and reward function estimators via a Bellman-type law of total variance to obtain tight bounds on the expected sum of the variances of value function estimates. This leads to a tighter dependence on the horizon in the function estimators. We also present numerical results to demonstrate the computational effectiveness of our proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。