用合成控制法评估大规模政策效果,提升精度并纠正机器学习偏差。
Post Launch Evaluation of Policies in a High-Dimensional Setting
- 先用近邻匹配选相似对照单元,再用高维学习估计反事实结果。
- 在六项大规模实验中,新方法显著提升政策效果估计准确率。
- 发现并修正机器学习带来的偏差,适合平台级产品评估者使用。
A/B测试是评估新政策、产品或决策影响的黄金标准,但成本高昂,可能使用户暴露于劣质选项。本文探讨在超大规模场景(达数亿单位)下,以“合成控制”方法替代传统A/B测试的可行性,尤其适用于仅部分单位受处理的情况。该方法利用未受影响单位的数据,估算处理单位的反事实结果,并与实际结果比较以衡量处理效应。核心挑战是插值偏差——当对照单位与处理单位差异大时易出现。为此,提出两阶段方法:第一阶段基于单位特征进行近邻匹配筛选相似对照;第二阶段采用适合高维数据的监督学习估计反事实。六项大规模实验验证了该方法能有效提升估计精度。然而分析发现,机器学习方法因权衡偏差与方差,会引入新偏差,影响结论可靠性。本文记录了此类偏差现象,并提出有效的去偏技术。
原文摘要 · Abstract (English)
A/B tests, also known as randomized controlled experiments (RCTs), are the gold standard for evaluating the impact of new policies, products, or decisions. However, these tests can be costly in terms of time and resources, potentially exposing users, customers, or other test subjects (units) to inferior options. This paper explores practical considerations in applying methodologies inspired by "synthetic control" as an alternative to traditional A/B testing in settings with very large numbers of units, involving up to hundreds of millions of units, which is common in modern applications such as e-commerce and ride-sharing platforms. This method is particularly valuable in settings where the treatment affects only a subset of units, leaving many units unaffected. In these scenarios, synthetic control methods leverage data from unaffected units to estimate counterfactual outcomes for treated units. After the treatment is implemented, these estimates can be compared to actual outcomes to measure the treatment effect. A key challenge in creating accurate counterfactual outcomes is interpolation bias, a well-documented phenomenon that occurs when control units differ significantly from treated units. To address this, we propose a two-phase approach: first using nearest neighbor matching based on unit covariates to select similar control units, then applying supervised learning methods suitable for high-dimensional data to estimate counterfactual outcomes. Testing using six large-scale experiments demonstrates that this approach successfully improves estimate accuracy. However, our analysis reveals that machine learning bias -- which arises from methods that trade off bias for variance reduction -- can impact results and affect conclusions about treatment effects. We document this bias in large-scale experimental settings and propose effective de-biasing techniques to address this challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。