从含隐变量的数据中自动学习工具变量表示,实现无偏因果推断。
Disentangled Representation Learning for Causal Inference with Instruments
- 基于变分自编码器学习可分离的工具变量表示
- 在合成与真实数据上优于现有因果估计方法
- 适合缺乏已知工具变量的因果分析场景
隐变量是基于观测数据进行因果推断的根本挑战。工具变量(IV)方法是一种实用的应对策略。现有基于IV的估计器需要已知的工具变量或强假设(如系统中存在两个及以上工具变量),限制了其应用范围。本文提出一种更宽松的假设:系统中存在一个工具变量代理,但不明确具体是哪个变量。我们提出一种基于变分自编码器(VAE)的解耦表征学习方法,从含有隐变量的数据中学习工具变量表示,并利用该表示对因果效应进行无偏估计。在合成数据和真实数据上的大量实验表明,所提算法优于现有的基于IV的估计器和基于VAE的估计器。
原文摘要 · Abstract (English)
Latent confounders are a fundamental challenge for inferring causal effects from observational data. The instrumental variable (IV) approach is a practical way to address this challenge. Existing IV based estimators need a known IV or other strong assumptions, such as the existence of two or more IVs in the system, which limits the application of the IV approach. In this paper, we consider a relaxed requirement, which assumes there is an IV proxy in the system without knowing which variable is the proxy. We propose a Variational AutoEncoder (VAE) based disentangled representation learning method to learn an IV representation from a dataset with latent confounders and then utilise the IV representation to obtain an unbiased estimation of the causal effect from the data. Extensive experiments on synthetic and real-world data have demonstrated that the proposed algorithm outperforms the existing IV based estimators and VAE-based estimators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。