arXiv:2502.12013cs.LGstat.ML2025-02

无监督下跨域生成反事实样本,突破传统方法依赖配对数据的限制。

Unsupervised Structural-Counterfactual Generation under Domain Shift

  • 通过区分因果变量的内在与域特异性成分,构建统一因果图框架。
  • 在合成数据上生成的反事实样本接近真实值,验证了方法有效性。
  • 适用于缺乏平行数据的跨域分析,如医疗、金融等场景的因果推断。

受跨域学习兴起的启发,我们提出一项新的生成建模挑战:基于源域的事实观测,在目标域中生成反事实样本。我们的方法在无监督范式下运行,无需平行或联合数据集,仅依赖各域的独立观测样本和因果图。该设定带来的挑战超越传统反事实生成。核心在于将外生原因解耦为效应内生与域内生两类,通过共享的效应内生外生变量,将各域的因果图整合为统一的联合因果图。我们在该联合框架中引入神经因果模型,实现标准可识别性假设下的准确反事实生成。此外,提出一种新型损失函数,在训练中有效分离效应内生与域内生变量。给定一个事实观测,框架结合源域效应内生变量的后验分布与目标域域内生变量的先验分布,合成所需反事实,遵循Pearl的因果层级。有趣的是,当域转移仅限于因果机制变化而无协变量偏移时,训练过程等价于求解条件最优传输问题。在合成数据上的实证评估表明,本框架生成的目标域反事实样本与真实值高度吻合。

原文摘要 · Abstract (English)

Motivated by the burgeoning interest in cross-domain learning, we present a novel generative modeling challenge: generating counterfactual samples in a target domain based on factual observations from a source domain. Our approach operates within an unsupervised paradigm devoid of parallel or joint datasets, relying exclusively on distinct observational samples and causal graphs for each domain. This setting presents challenges that surpass those of conventional counterfactual generation. Central to our methodology is the disambiguation of exogenous causes into effect-intrinsic and domain-intrinsic categories. This differentiation facilitates the integration of domain-specific causal graphs into a unified joint causal graph via shared effect-intrinsic exogenous variables. We propose leveraging Neural Causal models within this joint framework to enable accurate counterfactual generation under standard identifiability assumptions. Furthermore, we introduce a novel loss function that effectively segregates effect-intrinsic from domain-intrinsic variables during model training. Given a factual observation, our framework combines the posterior distribution of effect-intrinsic variables from the source domain with the prior distribution of domain-intrinsic variables from the target domain to synthesize the desired counterfactuals, adhering to Pearl's causal hierarchy. Intriguingly, when domain shifts are restricted to alterations in causal mechanisms without accompanying covariate shifts, our training regimen parallels the resolution of a conditional optimal transport problem. Empirical evaluations on a synthetic dataset show that our framework generates counterfactuals in the target domain that very closely resemble the ground truth.

反事实生成因果推断跨域学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。