用上下文学习估计因果效应,无需知道真实因果图。
Do-PFN: In-Context Learning for Causal Effect Estimation
- 在合成数据上预训练PFN,通过上下文学习从观测数据推断干预结果。
- 在多种因果结构下准确估计因果效应,无需依赖真实因果图。
- 适用于缺乏干预数据或因果知识的现实场景,对复杂因果结构稳健。
因果效应估计在多个科学领域至关重要。现有方法通常需要干预数据、真实的因果图知识,或依赖无混杂假设,限制了其在现实场景中的应用。在表格机器学习领域,先验数据拟合网络(PFNs)已在预训练后通过上下文学习实现最优预测性能。为检验该方法能否推广到更难的因果效应估计任务,我们基于多种因果结构(含干预)生成的合成数据预训练PFNs,使其在仅给定观测数据的情况下预测干预结果。在多个合成案例研究中,我们的方法无需了解底层因果图即可准确估计因果效应。消融实验进一步揭示了Do-PFN在不同因果特征数据集上的可扩展性与鲁棒性。
原文摘要 · Abstract (English)
Estimation of causal effects is critical to a range of scientific disciplines. Existing methods for this task either require interventional data, knowledge about the ground truth causal graph, or rely on assumptions such as unconfoundedness, restricting their applicability in real-world settings. In the domain of tabular machine learning, Prior-data fitted networks (PFNs) have achieved state-of-the-art predictive performance, having been pre-trained on synthetic data to solve tabular prediction problems via in-context learning. To assess whether this can be transferred to the harder problem of causal effect estimation, we pre-train PFNs on synthetic data drawn from a wide variety of causal structures, including interventions, to predict interventional outcomes given observational data. Through extensive experiments on synthetic case studies, we show that our approach allows for the accurate estimation of causal effects without knowledge of the underlying causal graph. We also perform ablation studies that elucidate Do-PFN's scalability and robustness across datasets with a variety of causal characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。