统一在线与离线实验的方差降低方法,揭示关键算法本质等价。
Unifying On- and Off-Policy Variance Reduction Methods
- 证明在线差异均值与离线逆倾向评分加最优控制变量数学等价。
- 回归调整法(如CUPED)与双重稳健估计在结构上完全一致。
- 帮助从业者理解并跨领域迁移实验优化策略。
连续高效的实验是网页用户应用实践成功的关键,涵盖在线A/B测试与离线策略评估。尽管两者目标相同——估算干预措施的增量价值,但常使用不同的术语和统计工具,彼此隔离。本文通过建立其典型方差降低方法的正式等价性,弥合这一鸿沟。我们证明:标准在线差异均值估计器在数学上等同于配备最优(方差最小化)加性控制变量的离线逆倾向评分估计器。进一步地,广泛使用的回归调整方法(如CUPED、CUPAC、ML-RATE)在结构上等同于双重稳健估计。这一统一看法深化了对常用方法的理解,并可指导在任一领域工作的研究人员与实践者。
原文摘要 · Abstract (English)
Continuous and efficient experimentation is key to the practical success of user-facing applications on the web, both through online A/B-tests and off-policy evaluation. Despite their shared objective -- estimating the incremental value of a treatment -- these domains often operate in isolation, utilising distinct terminologies and statistical toolkits. This paper bridges that divide by establishing a formal equivalence between their canonical variance reduction methods. We prove that the standard online Difference-in-Means estimator is mathematically identical to an off-policy Inverse Propensity Scoring estimator equipped with an optimal (variance-minimising) additive control variate. Extending this unification, we demonstrate that widespread regression adjustment methods (such as CUPED, CUPAC, and ML-RATE) are structurally equivalent to Doubly Robust estimation. This unified view extends our understanding of commonly used approaches, and can guide practitioners and researchers working on either class of problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。