厘清风险敏感强化学习中哪些效用函数可被高效学习
On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents
- 基于优化确定性等价的风险度量框架,分析样本复杂度
- 证明当效用函数定义域不完整时问题无法泛化学习
- 给出价值与策略学习的紧致下界,揭示关键参数影响
我们研究有限折扣马尔可夫决策过程中的风险敏感强化学习,假设存在生成模型。考虑一类称为优化确定性等价(OCE)的风险度量家族,包含熵风险、条件风险价值(CVaR)和均值-方差等重要度量。重点分析在递归OCE下的最优状态动作值函数(价值学习)和最优策略(策略学习)的样本复杂度。本文精确刻画了使对应OCE目标可实现概率近似正确(PAC)学习的效用函数$u$。分析了一种简单的基于模型的方法,并推导出PAC样本复杂度上界。证明当$u$的定义域$ ext{dom}(u)\neq \mathbb{R}$时,相应问题不可PAC学习。同时建立了价值与策略学习的下界,表明在状态-动作空间大小$SA$上的紧性;对于更受限的效用类,还得到依赖有效时长远距离$\frac{1}{1-γ}$的下界。具体地,对$\text{CVaR}_τ$,证明其对$τ$的正确依赖关系为$\frac{1}{τ^2}$,相比现有最优结果改进了$\frac{1}{τ}$倍,尽管其对$\frac{1}{1-γ}$的依赖为次优。
原文摘要 · Abstract (English)
We study risk-sensitive reinforcement learning in finite discounted MDPs, where a generative model of the MDP is assumed to be available. We consider a family or risk measures called the optimized certainty equivalent (OCE), which includes important risk measures such as entropic risk, CVaR, and mean-variance. Our focus is on the sample complexities of learning the optimal state-action value function (value learning) and an optimal policy (policy learning) under recursive OCE. We provide an exact characterization of utility functions $u$ for which the corresponding OCE defines an objective that is PAC-learnable. We analyze a simple model-based approach and derive PAC sample complexity bounds. We establish that whenever $u$ does not have full domain $\text{dom}(u)\neq \mathbb{R}$, the corresponding problem is not PAC-learnable. Finally, we establish corresponding lower bounds for both value and policy learning, demonstrating tightness in the size $SA$ of state-action space, and for a more restricted class of utilities, we derive lower bounds that makes the dependence on the effective horizon $\frac{1}{1-γ}$ explicit. Specifically, for $\text{CVaR}_τ$ we show that the correct dependence on $τ$ is $\frac{1}{τ^2}$, thus improving by a factor of $\frac{1}τ$ over state-of-the-art although our bound has a suboptimal dependence on $\frac{1}{1-γ}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。