保持变量关联性的表示学习,提升个体治疗效应估计精度
Structure Maintained Representation Learning Neural Network for Causal Inference
- 引入结构保持器,确保高维表示与原始特征间相关性不变
- 在真实医疗数据上,显著优于现有最优方法
- 适合需要精准个性化推断的医疗、社会科学领域
近年来,因果推断的关注点从平均治疗效应转向个体治疗效应。本文提出结构保持表示学习(SMRL)算法,通过引入结构保持器,在高维空间中维持基线协变量与其对应表示间的相关性,从而提升表示学习和对抗网络在个体治疗效应估计中的预测准确性。在表示层末端训练判别器,权衡表示平衡与信息损失;理论证明该判别器最小化了治疗效应估计误差的上界。通过考虑学习到的表示空间与原始协变量空间之间的相关性,有效解决分布平衡与信息丢失的权衡问题。在模拟数据和真实观测数据上进行了广泛实验,结果表明SMRL优于当前最先进方法。同时在MIMIC-III电子健康记录数据上验证了算法有效性。
原文摘要 · Abstract (English)
Recent developments in causal inference have greatly shifted the interest from estimating the average treatment effect to the individual treatment effect. In this article, we improve the predictive accuracy of representation learning and adversarial networks in estimating individual treatment effects by introducing a structure keeper which maintains the correlation between the baseline covariates and their corresponding representations in the high dimensional space. We train a discriminator at the end of representation layers to trade off representation balance and information loss. We show that the proposed discriminator minimizes an upper bound of the treatment estimation error. We can address the tradeoff between distribution balance and information loss by considering the correlations between the learned representation space and the original covariate feature space. We conduct extensive experiments with simulated and real-world observational data to show that our proposed Structure Maintained Representation Learning (SMRL) algorithm outperforms state-of-the-art methods. We also demonstrate the algorithms on real electronic health record data from the MIMIC-III database.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。