arXiv:2607.11508cs.LGcs.AI2026-07

构建通用因果发现模型,实现零样本结构推断。

CDFM: Towards a General-Purpose Causal Discovery Foundation Model

论文配图:CDFM: Towards a General-Purpose Causal Discovery Foundation Model
图 1 · 摘自论文原文
  • 将未知因果机制视为隐变量,设计可分解的变分框架。
  • 在海量合成数据上预训练,内化复杂统计不对称性。
  • 无需微调即可跨领域推理,适合大规模科学发现场景。

因果发现是从观测数据中恢复潜在因果结构的基础任务,在多个科学领域具有重要意义。过去几十年发展了大量算法,针对不同数据类型的因果机制设计专用流程,虽在多种应用中表现有效,但随着真实世界数据量与异质性持续增长,这种数据特定方法导致碎片化、依赖测试的范式,难以满足现代科学发现的需求。为此,我们提出因果发现基础模型(CDFM),作为统一的通用框架,实现零样本结构推断。为确保在未知领域可靠泛化,我们首先研究因果可识别性的理论边界,揭示因果先验机制的关键作用。基于此,构建一个原则性的变分框架,将未知因果机制视为隐变量,并将难以处理的边缘似然分解为若干可计算的学习模块。该变分分解为CDFM架构设计提供概念指导,同时结合全面的因果知识,大规模合成预训练数据。通过在高度多样化的合成结构因果模型空间上预训练,CDFM成功内化复杂统计不对称性。大量实验表明,CDFM持续优于传统算法,推动因果发现向通用基础模型范式转变。

原文摘要 · Abstract (English)

Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines. Over the past decades, numerous algorithms have been developed to tackle this challenge through workflows tailored to the specific causal mechanisms underlying each type of dataset, demonstrating effectiveness across a wide range of applications. However, as the volume and heterogeneity of real-world data continue to grow, this dataset-specific approach inevitably leads to a fragmented, test-driven paradigm that struggles to scale to the demands of modern scientific discovery. To address this, we formulate the Causal Discovery Foundation Model (CDFM) as a unified, general-purpose framework for zero-shot structural inference. To ensure reliable generalization across unknown domains, we first investigate the theoretical boundaries of causal identifiability, revealing the indispensable role of causal prior mechanisms in this process. Building on these insights, we formulate a principled variational framework that treats unknown causal mechanisms as latent variables and mathematically decomposes the intractable marginal likelihood into distinct, tractable learning modules. The variational decomposition provides a conceptual design principle for the architecture design of CDFM, while comprehensive causal knowledge guides the large-scale synthesis of our pretraining data. By pretraining on a massive, highly diverse space of synthetic structural causal models, CDFM successfully internalizes complex statistical asymmetries. Extensive experiments demonstrate that CDFM consistently outperforms traditional algorithms, driving a paradigm shift toward a general-purpose causal discovery foundation model.

因果发现基础模型变分推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。