对比了两种地理营销效果评估方法,发现机器学习方法更可靠。
Dynamic Synthetic Controls vs. Panel-Aware Double Machine Learning for Geo-Level Marketing Impact Estimation
- 用模拟器测试七种方法在五类复杂场景下的表现
- 传统合成控制法严重低估效果,置信区间几乎失效
- 面板双机器学习方法更稳健,适合真实商业场景
准确评估双边市场中地理级营销效果极具挑战:合成控制法(SCM)虽功效高,但系统性低估效应;而面板式双重机器学习(DML)极少与SCM进行对比。本文构建了一个开源、全文档化的模拟器,复现典型大规模地理推广:追踪N_unit个区域市场T_pre周预热期和T_post周的推广期,用户可调节所有关键参数,并在五种典型压力测试下评估两类方法:1)曲线基线趋势,2)异质响应延迟,3)处理组偏移冲击,4)非线性结果关联,5)控制组趋势漂移。共评估七种估计器:三种改进型增强合成控制(ASC)和四种面板式DML(TWFE、CRE/Mundlak、一阶差分、组内)。每种场景重复100次实验。结果显示,ASC在涉及非线性或外部冲击的复杂场景中表现出严重偏差,置信区间覆盖率接近零;而面板式DML显著降低偏差,恢复95%置信区间名义覆盖率,表现远更鲁棒。结论表明,尽管ASC提供简单基准,但在常见复杂情境下不可靠。因此提出‘先诊断再选择’框架:根据核心业务问题(如非线性趋势、响应延迟)匹配最适配的DML模型,为地理实验分析提供更可靠蓝图。
原文摘要 · Abstract (English)
Accurately quantifying geo-level marketing lift in two-sided marketplaces is challenging: the Synthetic Control Method (SCM) often exhibits high power yet systematically under-estimates effect size, while panel-style Double Machine Learning (DML) is seldom benchmarked against SCM. We build an open, fully documented simulator that mimics a typical large-scale geo roll-out: N_unit regional markets are tracked for T_pre weeks before launch and for a further T_post-week campaign window, allowing all key parameters to be varied by the user and probe both families under five stylized stress tests: 1) curved baseline trends, 2) heterogeneous response lags, 3) treated-biased shocks, 4) a non-linear outcome link, and 5) a drifting control group trend. Seven estimators are evaluated: three standard Augmented SCM (ASC) variants and four panel-DML flavors (TWFE, CRE/Mundlak, first-difference, and within-group). Across 100 replications per scenario, ASC models consistently demonstrate severe bias and near-zero coverage in challenging scenarios involving nonlinearities or external shocks. By contrast, panel-DML variants dramatically reduce this bias and restore nominal 95%-CI coverage, proving far more robust. The results indicate that while ASC provides a simple baseline, it is unreliable in common, complex situations. We therefore propose a 'diagnose-first' framework where practitioners first identify the primary business challenge (e.g., nonlinear trends, response lags) and then select the specific DML model best suited for that scenario, providing a more robust and reliable blueprint for analyzing geo-experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。