提出新型风险度量与多模式逼近方法,提升强化学习在有限时域下的风险控制能力。
Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation

- 引入小批量马尔可夫风险度量与多模式风险厌恶问题框架
- 理论证明高概率后悔界为 $\mathcal{O}(H^2 N^H \sqrt{K})$
- 适用于对风险敏感的短期决策场景,如任务分配与老虎机问题
针对风险规避型有限时域马尔可夫决策问题,本文提出一类特殊的马尔可夫一致风险度量——小批量风险度量,并定义了多模式风险厌恶问题类别,该类别推广了线性系统。基于此,我们设计了一种基于特征的 $Q$-learning 方法,采用多模式 $Q$-函数逼近,证明了高概率后悔界为 $\mathcal{O}(H^2 N^H \sqrt{K})$,其中 $H$ 为时域长度,$N$ 为小批量大小,$K$ 为博弈轮数。同时提出一种简化版算法,优化策略评估(反向)步骤。理论结果在随机指派问题与短时域多臂老虎机问题上得到验证。
原文摘要 · Abstract (English)
For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of multipattern risk-averse problems that generalizes the class of linear systems. We use both concepts in a feature-based $Q$-learning method with multipattern $Q$-factor approximation and we prove a high-probability regret bound of $\mathcal{O}\big(H^2 N^H \sqrt{ K}\big)$, where $H$ is the horizon, $N$ is the mini-batch size, and $K$ is the number of episodes. We also propose an economical version of the $Q$-learning method that streamlines the policy evaluation (backward) step. The theoretical results are illustrated on a stochastic assignment problem and a short-horizon multi-armed bandit problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。