arXiv:2603.03411stat.MLcs.LG2026-03

在未知软干预下实现可扩展的因果发现,提升结构恢复与泛化能力。

Scalable Contrastive Causal Discovery under Unknown Soft Interventions

  • 基于观测与干预数据对,通过对比跨域定向规则构建全局一致图
  • 在合成数据上实现更优结构恢复,且能处理未见图和新机制
  • 适用于大规模复杂系统,尤其适合干预目标未知的场景

观测因果发现仅能识别到马尔可夫等价类。虽然干预可减少歧义,但实际中干预常为软干预且目标未知。在许多现实场景中,仅观察到单一干预模式。本文提出一种可扩展的因果发现模型,适用于具有共享因果结构的成对观测与干预设置,且干预目标未知。该模型聚合子集级的PDAG,并应用对比跨域定向规则,在Meek闭包下构建全局一致的最大化PDAG,支持分布内与分布外泛化。理论上证明了模型在受限Ψ等价类下的正确性;进一步表明其渐近可恢复可识别的PDAG,且可比非对比方法定向更多边。合成数据实验显示,该模型在结构恢复、未见图泛化及大图可扩展性方面均有提升,消融实验验证了理论结果。

原文摘要 · Abstract (English)

Observational causal discovery is only identifiable up to the Markov equivalence class. While interventions can reduce this ambiguity, in practice interventions are often soft with multiple unknown targets. In many realistic scenarios, only a single intervention regime is observed. We propose a scalable causal discovery model for paired observational and interventional settings with shared underlying causal structure and unknown soft interventions. The model aggregates subset-level PDAGs and applies contrastive cross-regime orientation rules to construct a globally consistent maximal PDAG under Meek closure, enabling generalization to both in-distribution and out-of-distribution settings. Theoretically, we prove that our model is sound with respect to a restricted $Ψ$ equivalence class induced solely by the information available in the subset-restricted setting. We further show that the model asymptotically recovers the corresponding identifiable PDAG and can orient additional edges compared to non-contrastive subset-restricted methods. Experiments on synthetic data demonstrate improved causal structure recovery, generalization to unseen graphs with held-out causal mechanisms, and scalability to larger graphs, with ablations supporting the theoretical results.

因果发现软干预可扩展对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。