利用预训练提升治疗效果异质性估计精度,适用于医疗与教育决策。
Statistical Learning for Heterogeneous Treatment Effects: Pretraining, Prognosis, and Prediction
- 基于残差学习框架,融合预后预测与治疗效应估计的协同信号。
- 在高维协变量下显著降低误差与假阳性率,提升异质性检测能力。
- 适合需要个性化决策的领域,如精准医疗与教育政策制定。
稳健估计异质性治疗效应是个性化医疗到教育政策等领域的关键挑战。近年来,预测性机器学习为因果推断提供了灵活工具,但条件平均治疗效应(CATE)的准确估计仍面临困难,尤其在高维协变量下。本文提出一种预训练策略,利用现实应用中常见现象:对结果具有预后价值的因素也常能预测治疗效应异质性。例如,在医学中,同一生物信号通路的成分既影响基线风险又影响治疗反应。我们在此基础上,将该思想融入R-learner框架,通过残差化损失解决个体预测问题,并引入辅助信息以实现风险预测与因果效应估计间的跨任务学习。当此类协同关系存在时,该方法可更精确地探测信号,显著降低估计误差、减少假发现率,并提高检测异质性的统计功效。
原文摘要 · Abstract (English)
Robust estimation of heterogeneous treatment effects is a fundamental challenge for optimal decision-making in domains ranging from personalized medicine to educational policy. In recent years, predictive machine learning has emerged as a valuable toolbox for causal estimation, enabling more flexible effect estimation. However, accurately estimating conditional average treatment effects (CATE) remains a major challenge, particularly in the presence of many covariates. In this article, we propose pretraining strategies that leverage a phenomenon in real-world applications: factors that are prognostic of the outcome are frequently also predictive of treatment effect heterogeneity. In medicine, for example, components of the same biological signaling pathways frequently influence both baseline risk and treatment response. Specifically, we demonstrate our approach within the R-learner framework, which estimates the CATE by solving individual prediction problems based on a residualized loss. We use this structure to incorporate side information and develop models that can exploit synergies between risk prediction and causal effect estimation. In settings where these synergies are present, this cross-task learning enables more accurate signal detection, yields lower estimation error, reduced false discovery rates, and higher power for detecting heterogeneity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。