用深度学习提升因果推断中长期治疗效果估计的准确性
Deep Learning Methods for the Noniterative Conditional Expectation G-Formula for Causal Inference from Complex Observational Data
- 构建多任务循环网络统一估计时变变量联合条件分布
- 在复杂时间依赖场景下,深度学习模型偏差显著低于传统参数模型
- 适合处理具有复杂动态特征的观察性数据因果分析
g-公式可用于在一致性、正性与可交换性假设下,基于观察数据估计持续治疗策略的因果效应。非迭代条件期望(NICE)估计器同样需要准确估计时变治疗、混杂因素与结果的条件分布。传统上使用参数模型,但易因模型误设导致因果估计偏倚。本文提出一种统一的深度学习框架,采用多任务循环神经网络估计联合条件分布。通过模拟数据评估,发现当存在简单或复杂的时间依赖关系时,深度学习估计器对持续治疗策略在生存结局上的因果效应估计偏差更低,优于传统参数化NICE估计器。结果表明,该深度学习方法在复杂观察数据中估计长期治疗因果效应时,对模型误设更不敏感。
原文摘要 · Abstract (English)
The g-formula can be used to estimate causal effects of sustained treatment strategies using observational data under the identifying assumptions of consistency, positivity, and exchangeability. The non-iterative conditional expectation (NICE) estimator of the g-formula also requires correct estimation of the conditional distribution of the time-varying treatment, confounders, and outcome. Parametric models, which have been traditionally used for this purpose, are subject to model misspecification, which may result in biased causal estimates. Here, we propose a unified deep learning framework for the NICE g-formula estimator that uses multitask recurrent neural networks for estimation of the joint conditional distributions. Using simulated data, we evaluated our model's bias and compared it with that of the parametric g-formula estimator. We found lower bias in the estimates of the causal effect of sustained treatment strategies on a survival outcome when using the deep learning estimator compared with the parametric NICE estimator in settings with simple and complex temporal dependencies between covariates. These findings suggest that our Deep Learning g-formula estimator may be less sensitive to model misspecification than the classical parametric NICE estimator when estimating the causal effect of sustained treatment strategies from complex observational data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。