arXiv:2502.10883cs.LGcs.AI2025-02被引 2

提出新方法避免深度学习因果推断中的系统性偏差。

Learning Identifiable Structures Helps Avoid Bias in DNN-based Supervised Causal Learning

  • 通过预测骨架矩阵和v-结构张量,突破传统节点边架构局限。
  • 在合成与真实数据上显著优于现有DNN因果学习方法。
  • 适合需要可解释因果关系的医疗、金融等领域研究者。

因果发现是基于变量数据样本预测其因果关系的结构化预测任务。监督式因果学习(SCL)是该领域新兴范式。现有基于深度神经网络(DNN)的方法普遍采用“节点-边”架构:先为每个变量节点生成嵌入向量,再独立预测每条有向因果边。本文首次揭示该架构存在无法通过增大模型或数据规模消除的系统性偏差。为此,我们提出SiCL,一种基于DNN的SCL方法,同时预测骨架矩阵与v-张量(表示v-结构的三阶张量)。根据马尔可夫等价类(MEC)理论,骨架与v-结构在标准MEC设定下具有可识别性,因此其预测不受因果发现可识别性限制,使SiCL能够规避节点-边架构的系统偏差,并实现因果发现的一致估计器。此外,SiCL配备专门设计的成对编码模块,包含单向注意力层,用于建模节点对的内部与外部关系。在合成与真实世界基准上的实验结果表明,SiCL显著优于其他DNN-based SCL方法。

原文摘要 · Abstract (English)

Causal discovery is a structured prediction task that aims to predict causal relations among variables based on their data samples. Supervised Causal Learning (SCL) is an emerging paradigm in this field. Existing Deep Neural Network (DNN)-based methods commonly adopt the "Node-Edge approach", in which the model first computes an embedding vector for each variable-node, then uses these variable-wise representations to concurrently and independently predict for each directed causal-edge. In this paper, we first show that this architecture has some systematic bias that cannot be mitigated regardless of model size and data size. We then propose SiCL, a DNN-based SCL method that predicts a skeleton matrix together with a v-tensor (a third-order tensor representing the v-structures). According to the Markov Equivalence Class (MEC) theory, both the skeleton and the v-structures are identifiable causal structures under the canonical MEC setting, so predictions about skeleton and v-structures do not suffer from the identifiability limit in causal discovery, thus SiCL can avoid the systematic bias in Node-Edge architecture, and enable consistent estimators for causal discovery. Moreover, SiCL is also equipped with a specially designed pairwise encoder module with a unidirectional attention layer to model both internal and external relationships of pairs of nodes. Experimental results on both synthetic and real-world benchmarks show that SiCL significantly outperforms other DNN-based SCL approaches.

因果学习深度学习可识别性偏差避免

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。