将双机器学习用于面板数据,解决未观测异质性问题。
Double Machine Learning meets Panel Data -- Promises, Pitfalls, and Potential Solutions
- 用相关随机效应模型构建预测器,融入DML框架
- 样本量远大于观测混杂变量时估计准确
- 传统方法易受未观测异质性影响,新方法更稳健
使用机器学习算法估计因果效应可放松函数形式假设,但多数框架基于横截面数据,而研究者常拥有面板数据,传统方法中面板数据可用于处理个体间未观测异质性。本文探讨如何在存在未观测异质性的情况下将双/去偏机器学习(DML)适配至面板数据。该适配面临挑战:DML的交叉拟合程序依赖独立数据,且非线性可观测混杂下未观测异质性未必可加性分离。我们在多种模拟中评估多个直观方法的表现。尽管交叉拟合假设的违反对效果估计准确性影响较小,但多数方法未能充分处理未观测异质性。我们发现,在DML中使用基于相关随机效应方法(Mundlak, 1978)的预测模型,当样本量相对于可观测混杂变量数量较大时,可在各种设定下实现准确系数估计。同时,未观测异质性对可观测混杂变量的影响显著影响多数替代方法的表现。
原文摘要 · Abstract (English)
Estimating causal effect using machine learning (ML) algorithms can help to relax functional form assumptions if used within appropriate frameworks. However, most of these frameworks assume settings with cross-sectional data, whereas researchers often have access to panel data, which in traditional methods helps to deal with unobserved heterogeneity between units. In this paper, we explore how we can adapt double/debiased machine learning (DML) (Chernozhukov et al., 2018) for panel data in the presence of unobserved heterogeneity. This adaptation is challenging because DML's cross-fitting procedure assumes independent data and the unobserved heterogeneity is not necessarily additively separable in settings with nonlinear observed confounding. We assess the performance of several intuitively appealing estimators in a variety of simulations. While we find violations of the cross-fitting assumptions to be largely inconsequential for the accuracy of the effect estimates, many of the considered methods fail to adequately account for the presence of unobserved heterogeneity. However, we find that using predictive models based on the correlated random effects approach (Mundlak, 1978) within DML leads to accurate coefficient estimates across settings, given a sample size that is large relative to the number of observed confounders. We also show that the influence of the unobserved heterogeneity on the observed confounders plays a significant role for the performance of most alternative methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。