提出可同时保证高效与稳健的生成模型,用于从观测数据中估计潜在结果分布。
GDR-learners: Orthogonal Learning of Generative Models for Potential Outcomes
- 设计正交学习框架,使生成模型具备双重稳健性与准最优效率
- 在合成与半合成数据上表现优于现有方法,显著提升潜在结果分布估计精度
- 兼容多种主流生成模型,适用于因果推断与高维数据建模场景
已有深度生成模型用于从观测数据中估计潜在结果分布,但均缺乏奈曼正交性这一关键理论性质,因而不具备准最优效率和双重稳健性。本文提出一类通用的生成奈曼正交(双重稳健)学习器(GDR-learners),用于估计潜在结果的条件分布。所提方法具有灵活性,可基于多种前沿生成模型实现,包括:(a) 条件归一化流(GDR-CNFs)、(b) 条件生成对抗网络(GDR-CGANs)、(c) 条件变分自编码器(GDR-CVAEs)、(d) 条件扩散模型(GDR-CDMs)。与现有方法不同,本方法具备准最优效率与率双重稳健性,因而渐近最优。在一系列(半)合成实验中,GDR-learners 表现优异,显著优于现有方法,在潜在结果条件分布估计任务中取得更高准确性。
原文摘要 · Abstract (English)
Various deep generative models have been proposed to estimate potential outcomes distributions from observational data. However, none of them have the favorable theoretical property of general Neyman-orthogonality and, associated with it, quasi-oracle efficiency and double robustness. In this paper, we introduce a general suite of generative Neyman-orthogonal (doubly-robust) learners that estimate the conditional distributions of potential outcomes. Our proposed generative doubly-robust learners (GDR-learners) are flexible and can be instantiated with many state-of-the-art deep generative models. In particular, we develop GDR-learners based on (a) conditional normalizing flows (which we call GDR-CNFs), (b) conditional generative adversarial networks (GDR-CGANs), (c) conditional variational autoencoders (GDR-CVAEs), and (d) conditional diffusion models (GDR-CDMs). Unlike the existing methods, our GDR-learners possess the properties of quasi-oracle efficiency and rate double robustness, and are thus asymptotically optimal. In a series of (semi-)synthetic experiments, we demonstrate that our GDR-learners are very effective and outperform the existing methods in estimating the conditional distributions of potential outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。