arXiv:2412.19711stat.MLcs.LG2024-12被引 3

解决缺失结果数据下的治疗效果差异估计问题

Causal machine learning for heterogeneous treatment effects in the presence of missing outcome data

  • 提出mDR-learner和mEP-learner,用逆概率删失权重纠正样本偏差
  • 在模拟数据中表现优于传统方法,且满足正则效率条件
  • 适合处理医学研究中因缺失数据导致的子群体代表不足问题

在估计异质性治疗效应时,缺失结果数据会削弱某些人群的代表性,影响因果推断准确性。本文探讨了这一常被忽视的问题,分析了随机缺失(MAR)数据对条件平均治疗效应(CATE)估计器的影响。提出两种去偏机器学习估计器:mDR-learner 和 mEP-learner,分别将逆概率删失权重整合进 DR-learner 与 EP-learner,以缓解样本下采样问题。理论证明在合理假设下,这两类估计器具有原厂效率。通过模拟实验验证其优越性能,对比了现有 CATE 估计器及常用缺失数据处理方法。基于 GBSG2 临床试验数据,分析乳腺癌术后激素与非激素疗法的治疗效应异质性,并为实践者提供实施建议。

原文摘要 · Abstract (English)

When estimating heterogeneous treatment effects, missing outcome data can complicate treatment effect estimation, causing certain subgroups of the population to be poorly represented. In this work, we discuss this commonly overlooked problem and consider the impact that missing at random (MAR) outcome data has on causal machine learning estimators for the conditional average treatment effect (CATE). We propose two de-biased machine learning estimators for the CATE, the mDR-learner and mEP-learner, which address the issue of under-representation by integrating inverse probability of censoring weights into the DR-learner and EP-learner respectively. We show that under reasonable conditions, these estimators are oracle efficient, and illustrate their favorable performance through simulated data settings, comparing them to existing CATE estimators, including comparison to estimators which use common missing data techniques. We present an example of their application using the GBSG2 trial, exploring treatment effect heterogeneity when comparing hormonal therapies to non-hormonal therapies among breast cancer patients post surgery, and offer guidance on the decisions a practitioner must make when implementing these estimators.

因果推断缺失数据治疗效应机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。