提出新方法降低离策略评估的方差,提升评估准确性。
From Weighting to Modeling: A Nonparametric Estimator for Off-Policy Evaluation
- 用非参数模型构建权重,降低传统加权法的方差。
- 结合奖励预测进一步减少估计误差,性能优于现有方法。
- 适合需要高精度策略评估的研究者,尤其在数据分布不匹配时。
我们研究上下文博弈中的离策略评估问题,目标是利用历史数据(包含上下文、动作和回报)评估新策略。由于历史数据的动作分布与新策略不一致,传统逆概率加权(IPW)方法常因分母概率导致高方差。双重稳健(DR)虽通过建模回报降低方差,但未解决IPW本身的方差问题。本文提出非参数加权(NW)方法,利用非参数模型构造权重,实现低偏差且显著降低方差。为进一步减小方差,引入奖励预测机制,形成模型辅助非参数加权(MNW),通过显式建模和缓解奖励建模偏差,获得更准确的价值估计,无需保证标准双重稳健性质。大量实验证明,该方法在保持低偏差的同时,始终优于现有技术,显著降低价值估计方差。
原文摘要 · Abstract (English)
We study off-policy evaluation in the setting of contextual bandits, where we aim to evaluate a new policy using historical data that consists of contexts, actions and received rewards. This historical data typically does not faithfully represent action distribution of the new policy accurately. A common approach, inverse probability weighting (IPW), adjusts for these discrepancies in action distributions. However, this method often suffers from high variance due to the probability being in the denominator. The doubly robust (DR) estimator reduces variance through modeling reward but does not directly address variance from IPW. In this work, we address the limitation of IPW by proposing a Nonparametric Weighting (NW) approach that constructs weights using a nonparametric model. Our NW approach achieves low bias like IPW but typically exhibits significantly lower variance. To further reduce variance, we incorporate reward predictions -- similar to the DR technique -- resulting in the Model-assisted Nonparametric Weighting (MNW) approach. The MNW approach yields accurate value estimates by explicitly modeling and mitigating bias from reward modeling, without aiming to guarantee the standard doubly robust property. Extensive empirical comparisons show that our approaches consistently outperform existing techniques, achieving lower variance in value estimation while maintaining low bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。