解决目标人群测量不全混杂因子时的反事实预测偏差问题
Debiased Machine Learning for Conformal Prediction of Counterfactual Outcomes Under Runtime Confounding
- 基于半参数效率理论,构建去偏机器学习框架应对运行时混杂
- 仅需部分混杂变量测量即可实现准确覆盖率,收敛速度更快
- 适用于医疗决策等实际场景中混杂信息不完整的情况
数据驱动决策常依赖反事实结果预测。实践中,研究者通常在源数据集上训练反事实预测模型,以指导对可能不同的目标人群作出决策。近来,分布无关的置信区间方法(如分位数校准)被广泛用于生成目标人群中不同处理决策下反事实结果的预测区间。然而,现有方法要求所有用于源数据训练的治疗-结果关系混杂因素在目标人群中也被测量,若关键混杂因子未在目标人群测量,则可能导致覆盖不足。本文提出一种计算高效的去偏机器学习框架,可在仅部分混杂变量在目标人群中可测的情况下,依然保证预测区间的有效性,解决了常见的“运行时混杂”挑战。该方法基于半参数效率理论,证明其预测区间具有期望覆盖率且收敛速度优于标准方法。通过大量合成与半合成实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Data-driven decision making frequently relies on predicting counterfactual outcomes. In practice, researchers commonly train counterfactual prediction models on a source dataset to inform decisions on a possibly separate target population. Conformal prediction has arisen as a popular method for producing assumption-lean prediction intervals for counterfactual outcomes that would arise under different treatment decisions in the target population of interest. However, existing methods require that every confounding factor of the treatment-outcome relationship used for training on the source data is additionally measured in the target population, risking miscoverage if important confounders are unmeasured in the target population. In this paper, we introduce a computationally efficient debiased machine learning framework that allows for valid prediction intervals when only a subset of confounders is measured in the target population, a common challenge referred to as runtime confounding. Grounded in semiparametric efficiency theory, we show the resulting prediction intervals achieve desired coverage rates with faster convergence compared to standard methods. Through numerous synthetic and semi-synthetic experiments, we demonstrate the utility of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。