比较不同机器学习算法对双重机器学习置信区间的影响,发现模型选择显著影响推断可靠性。
Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence
- 对比解析法与自助法生成的置信区间,评估多种机器学习模型表现
- 样本量增大时,两类置信区间的覆盖概率反而下降,存在反直觉现象
- 实证发现农村程度越高,肥胖率越显著上升,且模型选择影响结果稳定性
双重机器学习(DML)是一种广泛用于处理效应估计的方法,可在保持有效推断的同时使用多种灵活的机器学习方法估计干扰参数。然而,实践中研究者需在众多机器学习算法中选择干扰模型,而该选择对DML方差估计的影响尚未充分阐明。本文通过全面模拟研究,比较了不同机器学习算法下DML置信区间的覆盖概率。我们对比了(1)基于DML理论推导的解析置信区间与(2)自助法置信区间,采用普通最小二乘、LASSO、随机森林、LightGBM和神经网络等模型,在不同数据生成设定下进行评估。通过偏差、置信区间宽度及最重要的覆盖概率来衡量性能。结果显示,解析与自助置信区间的覆盖性能在不同算法间存在显著差异,凸显模型选择对可靠推断的关键作用。令人意外的是,在许多情形下,随着样本量增加,解析与自助置信区间的覆盖概率反而下降。进一步利用美国县层级城乡差异的现实数据进行分析,发现(1)模型表现仍受算法选择影响,(2)更高程度的农村化对县层面肥胖率具有统计上显著的正向影响。
原文摘要 · Abstract (English)
Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference. In practice, however, applied researchers must choose among many machine learning algorithms for nuisance models, and the impact of this choice on the variance estimation of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different machine learning algorithms. In this study, we compare (1) analytical confidence intervals derived by DML theory versus (2) bootstrap confidence interval. We use a set of learners including ordinary least squares, LASSO, Random Forest, LightGBM, and Neural Networks under different data generation settings. We evaluate the performance across difference settings by bias, confidence interval width, and most importantly, coverage probability. Our results show substantial variability in coverage performance across analytical and bootstrap confidence intervals, highlighting that learner choice plays a critical role in reliable DML inference. Surprisingly, we find that in many settings, when sample size increases, the coverage probability of both DML analytical and bootstrap confidence interval decreases. We further investigate coverage probabilities using a real dataset on rural urban differences among U.S. counties. The real data analysis discovers that (1) the model performance still varies by the learner choices and (2) greater rurality has a statistically significant increasing effect on county level obesity prevalence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。