用在线学习优化自适应数据的因果效应估计,提升准确性。
Off-policy estimation with adaptively collected data: the power of online learning
- 通过在线学习动态调整估计,降低序列加权误差
- 在表格、线性及一般函数近似下均实现最优性能
- 适合做上下文博弈离线评估与因果推断的研究者
本文研究利用自适应收集数据估计处理效应的线性泛函问题,广泛应用于上下文赌博机的离线策略评估(OPE)和因果推断中的平均处理效应(ATE)估计。尽管某些增广逆倾向评分加权(AIPW)估计器具有渐近最优性质,但其在自适应数据下的非渐近理论仍不明确。为此,我们建立了AIPW估计器均方误差的通用上界,关键依赖于处理效应与其估计之间的序列加权误差。基于此,提出一种通用约简框架,通过在线学习生成一系列估计以最小化该误差。我们在三种情形下给出了具体实例:(1)表格情形;(2)线性函数逼近;(3)结果模型的一般函数逼近。进一步提供局部极小极大下界,证明使用无遗憾在线学习算法的AIPW估计器具有实例相关最优性。
原文摘要 · Abstract (English)
We consider estimation of a linear functional of the treatment effect using adaptively collected data. This task finds a variety of applications including the off-policy evaluation (\textsf{OPE}) in contextual bandits, and estimation of the average treatment effect (\textsf{ATE}) in causal inference. While a certain class of augmented inverse propensity weighting (\textsf{AIPW}) estimators enjoys desirable asymptotic properties including the semi-parametric efficiency, much less is known about their non-asymptotic theory with adaptively collected data. To fill in the gap, we first establish generic upper bounds on the mean-squared error of the class of AIPW estimators that crucially depends on a sequentially weighted error between the treatment effect and its estimates. Motivated by this, we also propose a general reduction scheme that allows one to produce a sequence of estimates for the treatment effect via online learning to minimize the sequentially weighted estimation error. To illustrate this, we provide three concrete instantiations in (\romannumeral 1) the tabular case; (\romannumeral 2) the case of linear function approximation; and (\romannumeral 3) the case of general function approximation for the outcome model. We then provide a local minimax lower bound to show the instance-dependent optimality of the \textsf{AIPW} estimator using no-regret online learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。