arXiv:2608.03868cs.LGcs.AI2026-08

提出可解释因果发现框架,每条边都有可审计的依据。

GENESIS: Towards Explainable Causal Discovery

论文配图:GENESIS: Towards Explainable Causal Discovery
图 1 · 摘自论文原文
  • 将图构建分解为可解释的决策点,结合统计证据与领域知识
  • 在所有数据量下实现100%决策可追溯性,优于传统方法
  • 适合需要透明决策过程的医疗、金融等高风险场景

从观测数据中进行因果发现面临两大挑战:纯统计方法在小样本情况下难以解决结构歧义;尽管大语言模型辅助的混合方法通过语义推理提升了结构恢复能力,但其推理过程对单个边的决策影响仍不透明。因此,现有混合方法无法满足基本要求——解释为何某个边被包含或排除在学习到的有向无环图(DAG)中。这在真实场景中至关重要,因不存在真值图,每个结构决策必须独立可验证。我们提出决策可追溯性这一形式化要求,即每个推断边需由可审计的统计证据、马尔可夫毯一致性或明确领域推理支持。我们提出GENESIS,一个可解释的混合因果发现框架,将图构建分解为可解释的决策点。首先识别并评分三节点结构模式(链、分叉、碰撞),建立透明的结构先验,然后逐步融合这些先验与观测证据,在统计证据不足时才引入领域知识。设计上,每个边决策均来自可审计的证据源。实验表明,GENESIS在所有设置下实现100%决策可追溯性,将可解释性作为首要目标。尽管有此额外要求,其在多数基准数据集上,所有样本条件下均优于纯统计因果发现方法(以结构汉明距离,SHD衡量),且性能接近当前最优的基于LLM的方法。

原文摘要 · Abstract (English)

Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely opaque. Consequently, existing hybrid methods fail to satisfy a fundamental requirement: explaining why a particular edge is included or excluded in the learned directed acyclic graph (DAG). This is critical in real-world applications, where no ground-truth DAG exists and every structural decision must be independently justified. We formalize this requirement as decision traceability, requiring every inferred edge to be supported by auditable statistical evidence, Markov Blanket consistency, or explicit domain reasoning. We propose GENESIS, an explainable hybrid CD framework that decomposes graph construction into interpretable decision points. GENESIS first identifies and scores three-node structural motifs, including chains, forks, and colliders, to establish transparent structural priors, then progressively refines the graph by integrating these priors with observational evidence, invoking domain knowledge only when statistical evidence is insufficient. By design, every edge decision is resolved through an auditable source of evidence. Experiments show that GENESIS achieves 100% decision traceability across all settings, establishing explainability as a first-class objective in causal discovery. Despite this additional requirement, GENESIS consistently outperforms purely statistical CD methods on the majority of benchmark datasets across all sample regimes in terms of Structural Hamming Distance (SHD), while achieving performance comparable to state-of-the-art LLM-assisted approaches.

因果发现可解释性大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。