针对概念漂移下的策略学习,提出更鲁棒的评估与优化方法。
Distributionally Robust Policy Learning under Concept Drifts
- 基于条件分布扰动构建双重稳健估计器
- 理论证明估计值渐近正态,且对低速估计也有效
- 算法在小样本下仍表现优异,适合高不确定性场景
分布鲁棒策略学习旨在找到在最坏分布偏移下表现良好的策略,但现有方法多考虑协变量与结果的联合分布最坏情况,可能过于保守。本文研究更精细的问题——仅协变量与结果的条件关系发生改变时的鲁棒策略学习。为此,我们提出一种双重稳健估计器,用于评估给定策略在一组扰动条件分布下的最坏平均回报。即使干扰参数以慢于根号n的速度估计,该策略价值估计量仍保持渐近正态性。进一步提出一个学习算法,在给定策略类Π内最大化估计策略价值,并证明其次优差距为κ(Π)·n⁻¹⁄²,其中κ(Π)是Π在汉明距离下的熵积分,n为样本量。匹配的下界表明该率最优。数值实验验证了方法显著优于现有基准。
原文摘要 · Abstract (English)
Distributionally robust policy learning aims to find a policy that performs well under the worst-case distributional shift, and yet most existing methods for robust policy learning consider the worst-case joint distribution of the covariate and the outcome. The joint-modeling strategy can be unnecessarily conservative when we have more information on the source of distributional shifts. This paper studies a more nuanced problem -- robust policy learning under the concept drift, when only the conditional relationship between the outcome and the covariate changes. To this end, we first provide a doubly-robust estimator for evaluating the worst-case average reward of a given policy under a set of perturbed conditional distributions. We show that the policy value estimator enjoys asymptotic normality even if the nuisance parameters are estimated with a slower-than-root-$n$ rate. We then propose a learning algorithm that outputs the policy maximizing the estimated policy value within a given policy class $Π$, and show that the sub-optimality gap of the proposed algorithm is of the order $κ(Π)n^{-1/2}$, where $κ(Π)$ is the entropy integral of $Π$ under the Hamming distance and $n$ is the sample size. A matching lower bound is provided to show the optimality of the rate. The proposed methods are implemented and evaluated in numerical studies, demonstrating substantial improvement compared with existing benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。