在数据丰富的场景下,提出一种能处理隐藏混杂因素的因果推断框架。
A Causal Inference Framework for Data Rich Environments
- 融合结构因果模型与潜在结果模型,构建统一框架。
- 证明了平均处理效应等关键参数可被识别并一致估计。
- 适用于高维面板数据,适合做政策评估的研究者使用。
我们为存在未观测混杂因素的“数据丰富”场景(即单位数量多、每单位测量维度高)提出了一个正式的反事实估计模型。该模型连接了图形模型文献中的结构因果模型与潜在结果文献中的潜在因子模型视角。我们展示了经典潜在结果与处理分配模型如何嵌入本框架。给出了平均处理效应、处理组平均处理效应及非处理组平均处理效应的可识别性论证。对于任意对某一扰动参数具有足够快估计误差率的估计器,我们证明其对这些因果参数具有一致性。随后表明主成分回归是一种满足条件的估计器,并分析了潜在结果函数所需最小光滑度以保证一致性。
原文摘要 · Abstract (English)
We propose a formal model for counterfactual estimation with unobserved confounding in "data-rich" settings, i.e., where there are a large number of units and a large number of measurements per unit. Our model provides a bridge between the structural causal model view of causal inference common in the graphical models literature with that of the latent factor model view common in the potential outcomes literature. We show how classic models for potential outcomes and treatment assignments fit within our framework. We provide an identification argument for the average treatment effect, the average treatment effect on the treated, and the average treatment effect on the untreated. For any estimator that has a fast enough estimation error rate for a certain nuisance parameter, we establish it is consistent for these various causal parameters. We then show principal component regression is one such estimator that leads to consistent estimation, and we analyze the minimal smoothness required of the potential outcomes function for consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。