基于数据构建鲁棒强化学习模型,确保小样本下性能可靠
Data-driven robust Markov decision processes on Borel spaces: performance guarantees via an axiomatic approach
- 用距离函数定义不确定性集,从数据中自适应构造鲁棒策略
- 样本量增大时,鲁棒解与真实最优解一致,小样本下保证上界
- 适用于数据有限、对安全性要求高的强化学习场景
针对扰动分布未知的马尔可夫决策过程(MDP),本文采用鲁棒马尔可夫决策过程(RMDP)方法。通过构造未知扰动分布的样本分布,并将不确定性集定义为以该样本分布为中心、距离函数值低于某阈值的集合,结合弱收敛与距离函数收敛的关系,证明了在样本量增加时,鲁棒最优值函数和离样本值函数均一致收敛至真实最优值函数。对于有限样本,证明了鲁棒最优值函数以高概率作为离样本值函数的上界,并给出了概率收敛速率、样本复杂度及分布外性能边界。这些结果依赖于距离函数满足某种浓度不等式,文献中多个常用距离函数均满足此条件。此外,对比分析表明,传统经验MDP无法满足此类有限样本保证。
原文摘要 · Abstract (English)
We consider Markov decision processes (MDPs) with unknown disturbance distribution and address this problem using the robust Markov decision process (RMDP) approach. We construct the empirical distribution of the unknown disturbance distribution and characterize our ambiguity set of distributions as the sublevel set of a nonnegative distance function from the empirical distribution. By connecting the weak convergence of distributions to convergence with respect to the distance function, we prove that the robust optimal value function and the out-of-sample value function converge to the true optimal value function with increasing sample-sizes. We establish that, for finite sample-sizes, the robust optimal value function serves as a high probability upper bound on the out-of-sample value function. We also obtain probabilistic convergence rates, sample complexity bounds, and out-of-distribution performance bounds. The finite sample performance guarantees rely on the distance function satisfying a certain concentration type inequality. Several well-studied distances in the literature meet the requirements imposed on the distance function. We also analyze the data-driven properties of empirical MDPs and demonstrate that, unlike our data-driven RMDPs, empirical MDPs fail to satisfy some of the finite sample performance guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。