arXiv:2603.20842cs.LG2026-03被引 1

用弱先验知识提升因果发现模型的实用性与鲁棒性

A Knowledge-Informed Pretrained Model for Causal Discovery

  • 通过双源编码器-解码器结构融合观测数据与弱先验知识
  • 在多种数据分布下优于现有基线,对不同图密度和变量规模均表现稳定
  • 适合缺乏强干预数据但有领域常识的科研与工业场景

因果发现虽被广泛研究,但多数方法依赖强假设,或过度依赖昂贵的干预信号与部分真实标签作为强先验,或完全数据驱动而缺乏指导,限制了实际应用。针对现实中仅具备粗粒度领域知识的情况,我们提出一种知识引导的预训练因果发现模型,将弱先验作为合理中间路径。模型采用双源编码器-解码器架构,以知识感知方式处理观测数据。设计多样化预训练数据集与课程学习策略,使模型能平滑适应不同先验强度、图密度及变量规模下的机制变化。在分布内、分布外及真实数据集上的大量实验表明,该模型持续优于现有基线,具备强鲁棒性与实用价值。

原文摘要 · Abstract (English)

Causal discovery has been widely studied, yet many existing methods rely on strong assumptions or fall into two extremes: either depending on costly interventional signals or partial ground truth as strong priors, or adopting purely data driven paradigms with limited guidance, which hinders practical deployment. Motivated by real-world scenarios where only coarse domain knowledge is available, we propose a knowledge-informed pretrained model for causal discovery that integrates weak prior knowledge as a principled middle ground. Our model adopts a dual source encoder-decoder architecture to process observational data in a knowledge-informed way. We design a diverse pretraining dataset and a curriculum learning strategy that smoothly adapts the model to varying prior strengths across mechanisms, graph densities, and variable scales. Extensive experiments on in-distribution, out-of distribution, and real-world datasets demonstrate consistent improvements over existing baselines, with strong robustness and practical applicability.

因果发现预训练模型知识引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。