新数据集+基准测试揭示治疗效应估计的不一致性,双稳健方法显著更优。
Benchmarking Estimators for Natural Experiments: A Novel Dataset and a Doubly Robust Algorithm
- 构建自然实验新数据集与合成结果基准,系统评估20余种估计器
- 双稳健方法在多种现实条件下表现远超复杂模型,误差低一个数量级
- 开源工具包支持新数据与算法扩展,适合因果推断研究者使用
从早期儿童阅读非营利组织获取新型自然实验数据集。令人意外的是,对同一数据集应用超过20种已有估计器,得出的非营利组织效果评估结果严重不一致。为解决此问题,我们基于领域专家指导设计合成结果,建立用于评估估计器准确性的基准,全面考察样本量、处理相关性及倾向得分精度等真实条件下的表现。基于该基准发现:基于简单回归调整的双稳健估计器,普遍比其他更复杂的估计器性能高出一个数量级。为深化对双稳健估计器的理解,我们推导出使用数据分割获得无偏估计时任意此类估计器方差的闭式表达式,据此提出一种新型双稳健估计器,其在回归调整中采用新颖损失函数。数据集与基准已以模块化Python包形式发布,便于新数据集和新估计器的集成。
原文摘要 · Abstract (English)
Estimating the effect of treatments from natural experiments, where treatments are pre-assigned, is an important and well-studied problem. We introduce a novel natural experiment dataset obtained from an early childhood literacy nonprofit. Surprisingly, applying over 20 established estimators to the dataset produces inconsistent results in evaluating the nonprofit's efficacy. To address this, we create a benchmark to evaluate estimator accuracy using synthetic outcomes, whose design was guided by domain experts. The benchmark extensively explores performance as real world conditions like sample size, treatment correlation, and propensity score accuracy vary. Based on our benchmark, we observe that the class of doubly robust treatment effect estimators, which are based on simple and intuitive regression adjustment, generally outperform other more complicated estimators by orders of magnitude. To better support our theoretical understanding of doubly robust estimators, we derive a closed form expression for the variance of any such estimator that uses dataset splitting to obtain an unbiased estimate. This expression motivates the design of a new doubly robust estimator that uses a novel loss function when fitting functions for regression adjustment. We release the dataset and benchmark in a Python package; the package is built in a modular way to facilitate new datasets and estimators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。