arXiv:2506.10677stat.MLcs.LG2025-06KDD被引 1

利用系统相似性提升A/B测试精度,更准更快得出结论

Exploiting Similarities in A/B Testing with Off-Policy Estimation

  • 用离策略估计法捕捉新旧系统决策相似性
  • 在系统相似时误差显著降低,最坏情况下不差于传统方法
  • 适合有结构相似性的在线决策优化场景

我们研究A/B测试,即评估新决策系统相对于基线系统性能提升的标准方法。传统A/B测试将两个系统视为黑箱,忽略了它们之间可能存在的相似性。实践中,新系统与基线系统通常并非本质不同,常具有显著结构相似性,这种相似性可通过其做出相似决策的倾向性体现。我们发现,尽管常用的均值差异估计器无偏,但在系统相似时统计效率低下。通过引入离策略估计框架,我们提出一类可利用系统倾向性以改善集中性质的A/B测试估计器。该族估计器灵活,可适配实际决策场景;方法简单、对倾向性误设稳健,在系统具相似性时显著更准确,且在无相似性时自动退化为均值差异估计器。理论分析与实证研究均验证了其高效性与实用性。

原文摘要 · Abstract (English)

We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline. Traditional A/B testing treats both systems as black boxes, ignoring potential similarities between them. In practice, however, new and baseline systems are rarely radically different and often share significant structure, which can be captured by their propensities to make similar decisions. We show that in such cases, the commonly used difference-in-means estimator, though unbiased, is statistically suboptimal. Leveraging off-policy estimation, we introduce a family of A/B testing estimators that exploit the propensities of the tested systems to achieve improved concentration properties. This family is flexible enough to be tailored to practical decision-making. The resulting estimators are simple, robust to propensities misspecification, substantially more accurate when the tested systems exhibit similarities, and gracefully fall back to the difference-in-means estimator when such similarities are absent. Our theoretical analysis and empirical studies confirm their efficiency and practicality.

A/B测试离策略估计决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。