无需理想假设,利用数据异质性识别潜变量,估计长期个体因果效应。
Long-Term Individual Causal Effect Estimation via Identifiable Latent Representation Learning
- 利用多源数据天然异质性识别潜混杂因子
- 理论证明潜变量可识别,实现长期因果效应定位
- 在合成与半合成数据上验证了方法有效性
结合长期观测数据与短期实验数据来估计长期个体因果效应,在诸多实际场景中至关重要但极具挑战。现有方法依赖于理想化假设,如潜混杂无偏假设或等量混杂偏移假设,但在真实应用中这些假设常被违反,限制了其实际效果。本文提出一种无需上述假设的长期个体因果效应估计方法。具体而言,我们利用数据的自然异质性(如多源数据)来识别潜混杂因子,显著降低对理想假设的依赖。实践中,设计了一种基于潜表示学习的长期因果效应估计算法;理论上,建立了潜混杂因子的可识别性,并进一步实现长期效应的可识别性。在多个合成与半合成数据集上的大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Estimating long-term causal effects by combining long-term observational and short-term experimental data is a crucial but challenging problem in many real-world scenarios. In existing methods, several ideal assumptions, e.g. latent unconfoundedness assumption or additive equi-confounding bias assumption, are proposed to address the latent confounder problem raised by the observational data. However, in real-world applications, these assumptions are typically violated which limits their practical effectiveness. In this paper, we tackle the problem of estimating the long-term individual causal effects without the aforementioned assumptions. Specifically, we propose to utilize the natural heterogeneity of data, such as data from multiple sources, to identify latent confounders, thereby significantly avoiding reliance on idealized assumptions. Practically, we devise a latent representation learning-based estimator of long-term causal effects. Theoretically, we establish the identifiability of latent confounders, with which we further achieve long-term effect identification. Extensive experimental studies, conducted on multiple synthetic and semi-synthetic datasets, demonstrate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。