arXiv:2504.14937cs.LGcs.DB2025-04被引 2

提出可保留因果信息的图摘要方法,提升高维数据下因果推断的可解释性与鲁棒性。

Causal DAG Summarization (Full Version)

  • 设计兼顾简化与因果信息保留的图摘要目标函数
  • 在6个真实数据集上验证,摘要图可直接用于可靠推断
  • 适合需要可解释因果分析的高维数据研究者

因果推断帮助研究人员发现因果关系,实现科学洞察。准确的因果估计需识别混杂变量以避免错误发现。Pearl的因果模型使用因果DAG识别混杂变量,但错误的DAG会导致不可靠的因果结论。然而,在高维数据中,因果DAG通常过于复杂,超出人类可验证范围。图摘要成为合理下一步,但现有通用图摘要方法不适用于因果DAG。本文提出一种因果图摘要目标,平衡图简化以提升可理解性,同时保留关键因果信息以确保可靠推断。我们开发了一种高效的贪心算法,并证明摘要因果DAG可直接用于推断,且对假设误设更鲁棒,提升了因果推断的稳健性。在六个真实数据集上与三种现有方法对比,结果表明该算法能有效处理高维数据,生成既保障可靠因果推断又具备抗误设能力的摘要图。

原文摘要 · Abstract (English)

Causal inference aids researchers in discovering cause-and-effect relationships, leading to scientific insights. Accurate causal estimation requires identifying confounding variables to avoid false discoveries. Pearl's causal model uses causal DAGs to identify confounding variables, but incorrect DAGs can lead to unreliable causal conclusions. However, for high dimensional data, the causal DAGs are often complex beyond human verifiability. Graph summarization is a logical next step, but current methods for general-purpose graph summarization are inadequate for causal DAG summarization. This paper addresses these challenges by proposing a causal graph summarization objective that balances graph simplification for better understanding while retaining essential causal information for reliable inference. We develop an efficient greedy algorithm and show that summary causal DAGs can be directly used for inference and are more robust to misspecification of assumptions, enhancing robustness for causal inference. Experimenting with six real-life datasets, we compared our algorithm to three existing solutions, showing its effectiveness in handling high-dimensional data and its ability to generate summary DAGs that ensure both reliable causal inference and robustness against misspecifications.

因果推断图摘要DAG建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。