通过利用策略重叠降低实验方差,加速A/B测试。
Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap

- 将随机分配视为元策略,用反事实估计减少噪声。
- 方差随策略差异缩小,而非依赖原始结果方差。
- 适合评估推荐系统、大模型接口等高成本场景。
在线控制实验是在线平台假设检验的金标准。尽管广泛应用,其运行成本高昂,方差问题严重制约了对处理效应的统计功效。传统方差缩减方法依赖模型控制变量来降低结果噪声,但对竞争策略间的结构性关系视而不见。本文发现标准A/B测试协议存在关键低效:当处理与对照策略在某动作上一致时,对应结果仅贡献噪声,不提供处理效应信号,从而无谓扩大置信区间。为此,提出一种新实验协议,利用策略重叠加速实验。核心思路是将随机处理分配机制视为元策略,采用Δ-离策略估计方法获取平均处理效应的无偏估计。理论上证明该方法在一般情况下等同于标准A/B测试,但其方差随策略间差异变化,而非原始结果方差。因此,当策略存在共同支持域时,性能优于标准均值差估计器,且当重叠区域存在非零残差方差时改进严格成立。实验证实了这些理论洞察,对推荐系统、信息检索管道及大语言模型界面的现实评估具有重要潜力。
原文摘要 · Abstract (English)
Online controlled experiments are the gold standard for hypothesis testing in online platforms. Notwithstanding their ubiquity, they are notoriously expensive to run, and issues of variance hamper statistical power in assessing treatment effects. While standard variance reduction techniques leverage model-based control variates to reduce outcome noise, they remain agnostic to potential structural relationships between competing policies. In this work, we identify a critical inefficiency in the standard A/B-testing protocol: when a treatment and control policy agree on an action, the resulting outcome contributes noise but no signal regarding the treatment effect -- unnecessarily inflating confidence intervals. We propose a novel experimental protocol that exploits this policy overlap to accelerate experimentation. The key insight is to frame the randomised treatment assignment mechanism as a meta-policy, and leverage $Δ$-Off-Policy Estimation methods to obtain unbiased estimates for average treatment effects. We prove analytically that our approach recovers standard A/B-testing practices in the general case, but that its variance scales with the divergence between policies rather than raw outcome variance. Hence, we dominate the standard Difference-in-Means estimator whenever policies have common support, and the improvement is strict whenever the overlap region contributes non-zero residual variance. Empirical results corroborate these theoretical insights -- holding promise for significant impact on the real-world evaluation of recommender systems, information retrieval pipelines, and large language model interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。