arXiv:2507.18118stat.MLcs.LG2025-07被引 3

提出双臂老虎机框架,提升A/B测试统计功效。

A Two-armed Bandit Framework for A/B Testing

  • 用双重稳健估计生成伪结果,增强评估可靠性。
  • 结合双臂老虎机构造检验统计量,提升检测灵敏度。
  • 适合需要高精度决策的技术公司和数据科学家。

A/B测试被现代科技公司广泛用于策略评估与产品部署,旨在比较新策略与基准控制下的效果。文献中多种因果推断和强化学习方法适用于A/B测试。本文提出一种双臂老虎机框架,旨在提升现有方法的统计功效。该方法包含三个步骤:(i) 使用双重稳健估计生成伪结果;(ii) 利用双臂老虎机框架构建检验统计量;(iii) 采用置换法计算p值。通过渐近理论、数值实验及某网约车公司的实际数据验证了该方法的有效性,结果表明其在性能上优于现有方法。

原文摘要 · Abstract (English)

A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy against a standard control. Various causal inference and reinforcement learning methods developed in the literature are applicable to A/B testing. This paper introduces a two-armed bandit framework designed to improve the power of existing approaches. The proposed procedure consists of three main steps: (i) employing doubly robust estimation to generate pseudo-outcomes, (ii) utilizing a two-armed bandit framework to construct the test statistic, and (iii) applying a permutation-based method to compute the $p$-value. We demonstrate the efficacy of the proposed method through asymptotic theories, numerical experiments and real-world data from a ridesharing company, showing its superior performance in comparison to existing methods.

A/B测试因果推断统计检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。