用模拟器辅助学习因果表示,提升后处理协变量下的个体治疗效应估计。
Leveraging a Simulator for Learning Causal Representations from Post-Treatment Covariates for CATE
- 通过模拟器生成反事实监督信号,学习与治疗无关的因果表征。
- 理论证明真实与模拟数据分布差异越大,CATE误差越高,且可量化。
- 提出SimPONet方法,自动调节模拟器对学习目标的影响,适合复杂现实场景。
治疗效应估计旨在评估不同干预对个体结果的影响。现有方法基于观测数据估计条件平均治疗效应(CATE),依赖于治疗前协变量与结果的时序关系,并假设满足正性与无混杂性。本文研究一种新场景:协变量和结果均在治疗后采集,此时传统方法失效。我们指出,后处理协变量导致CATE不可识别,恢复其可识别性需学习与治疗无关的因果表示。先前工作表明,在存在反事实监督的前提下,可通过对比学习实现该目标。但反事实数据稀少,已有研究尝试使用模拟器提供合成反事实监督。本文系统分析模拟器在CATE估计中的作用:对比多个基线方法并揭示其局限性;建立泛化误差上界,量化联合训练真实与模拟分布时的真实-模拟差异对CATE误差的影响;进而提出新方法SimPONet,其损失函数源自上述上界,并能根据模拟器与任务的相关性动态调整其影响权重。我们在多种数据生成过程(DGPs)下进行实验,系统地改变真实与模拟分布差距,验证SimPONet相较于当前最优基线的有效性。
原文摘要 · Abstract (English)
Treatment effect estimation involves assessing the impact of different treatments on individual outcomes. Current methods estimate Conditional Average Treatment Effect (CATE) using observational datasets where covariates are collected before treatment assignment and outcomes are observed afterward, under assumptions like positivity and unconfoundedness. In this paper, we address a scenario where both covariates and outcomes are gathered after treatment. We show that post-treatment covariates render CATE unidentifiable, and recovering CATE requires learning treatment-independent causal representations. Prior work shows that such representations can be learned through contrastive learning if counterfactual supervision is available in observational data. However, since counterfactuals are rare, other works have explored using simulators that offer synthetic counterfactual supervision. Our goal in this paper is to systematically analyze the role of simulators in estimating CATE. We analyze the CATE error of several baselines and highlight their limitations. We then establish a generalization bound that characterizes the CATE error from jointly training on real and simulated distributions, as a function of the real-simulator mismatch. Finally, we introduce SimPONet, a novel method whose loss function is inspired from our generalization bound. We further show how SimPONet adjusts the simulator's influence on the learning objective based on the simulator's relevance to the CATE task. We experiment with various DGPs, by systematically varying the real-simulator distribution gap to evaluate SimPONet's efficacy against state-of-the-art CATE baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。