用历史依赖策略估计降低离线评估误差,理论解释了为何更准
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
- 通过偏差-方差分解,揭示历史依赖策略可降低重要性采样方差
- 估计策略考虑更长历史时,渐近方差持续下降,有限样本偏差上升
- 适用于多种离线评估方法,涵盖参数与非参数估计场景
本文研究强化学习中基于重要性采样的离线策略评估(OPE),重点分析行为策略估计对性能的影响。已有实验证明,即使真实行为策略为马尔可夫性,使用历史依赖策略估计仍能降低均方误差(MSE)。本文首次从理论上揭示这一悖论:普通重要性采样(IS)估计器的MSE可分解为偏差与方差,历史依赖策略估计虽增加有限样本偏差,但显著降低渐近方差。随着估计策略所依赖的历史长度增加,方差持续下降。该结论进一步扩展至序列化IS、双重稳健及边际化IS等主流估计算法,无论采用参数或非参数方式估计行为策略均成立。
原文摘要 · Abstract (English)
This paper studies off-policy evaluation (OPE) in reinforcement learning with a focus on behavior policy estimation for importance sampling. Prior work has shown empirically that estimating a history-dependent behavior policy can lead to lower mean squared error (MSE) even when the true behavior policy is Markovian. However, the question of why the use of history should lower MSE remains open. In this paper, we theoretically demystify this paradox by deriving a bias-variance decomposition of the MSE of ordinary importance sampling (IS) estimators, demonstrating that history-dependent behavior policy estimation decreases their asymptotic variances while increasing their finite-sample biases. Additionally, as the estimated behavior policy conditions on a longer history, we show a consistent decrease in variance. We extend these findings to a range of other OPE estimators, including the sequential IS estimator, the doubly robust estimator and the marginalized IS estimator, with the behavior policy estimated either parametrically or non-parametrically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。