arXiv:2606.21806cs.LG2026-06

提出新模型生成受混杂因素影响的图像,能区分因果效应。

Causal Variational Deep Embedding: A Family of Interventional Generators for Confounded Images

论文配图:Causal Variational Deep Embedding: A Family of Interventional Generators for Confounded Images
图 1 · 摘自论文原文
  • 用离散潜变量建模隐含混杂因子,分离出可控因果机制。
  • 在多个图像数据集上生成多样干预样本,FID指标优于无混杂基准。
  • 适合需要可解释生成与因果推断的研究者使用。

深度生成模型复现训练数据的观测分布,继承其中的虚假关联。常见来源是未观测到的混杂因子,它同时影响用户希望在采样时控制的属性和预期会变化的属性。现有因果生成方法通过强结构假设确定单一干预分布;但在图像领域,此类假设往往不成立,数据通常符合一组不同的因果机制——即可行的干预分布区域。我们提出 CauVaDE(因果变分深度嵌入),基于一个标准扩展结构因果模型(SCM),其中未观测混杂因子以有界支持的离散潜簇形式存在,连续变化则由独立噪声吸收。证明该标准类在观测与干预的Wasserstein距离下稠密覆盖与给定因果图兼容的所有扩展SCM,并实例化为混合变分自编码器,其簇变量充当标准混杂因子。对簇后验施加权重为γ的熵正则项,从而追踪一系列拟合观测数据且具有可比似然性的候选因果效应,覆盖可行区域。在图像数据基准上的实验表明,CauVaDE 能生成多样化的干预样本,并在FID上优于无混杂参考模型。

原文摘要 · Abstract (English)

Deep generative models reproduce the observational distribution of their training data, inheriting any spurious associations it contains. A common source is an unobserved confounder that shapes both an attribute the user wants to control at sampling time and an attribute expected to vary in response. Existing causal generative approaches resolve the resulting ambiguity by imposing structural assumptions strong enough to single out one interventional distribution; in image domains, such assumptions are rarely warranted, and the data is generally consistent with a set of distinct causal mechanisms -- a feasible region of interventional distributions. We propose CauVaDE (Causal Variational Deep Embedding), built on a canonical augmented SCM in which the unobserved confounder collapses, without loss of generality, into a discrete latent cluster of bounded support while continuous variation is absorbed into independent noises. We prove that this canonical class is dense, in both observational and interventional Wasserstein distance, in the class of augmented SCMs compatible with a given causal diagram, and instantiate it as a mixture variational autoencoder whose cluster variable plays the role of the canonical confounder. An entropy regularizer with weight $γ$ on the cluster posterior then traces a family of candidate causal effects that fit the observational data to comparable likelihood while spanning the feasible region. Experiments on image data benchmarks show that CauVaDE produces diverse interventional samples and improves FID against an unconfounded reference.

因果生成变分自编码器图像生成混杂因子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。