arXiv:2604.08890cs.LGcs.AI2026-04

提出基于最小单元的图因果建模方法,保障因果有效性。

A Closer Look at the Application of Causal Inference in Graph Representation Learning

  • 以图数据最小不可分单元为基础构建因果模型
  • 证明聚合操作会破坏因果推断假设,导致无效结论
  • 可无缝集成到现有图学习流程,适合关注因果可靠性的研究者

图表示学习中的因果关系建模仍是基础性挑战。现有方法常借鉴因果推断理论识别因果子图或缓解混杂因素,但由于图结构数据固有的复杂性,这些方法常将多种图元素合并为单一因果变量,这一操作可能违反因果推断的核心假设。本文证明此类聚合会损害因果有效性。基于此,我们提出一个以图数据最小不可分单元为基础的理论模型,确保因果有效性。在此基础上,我们分析了实现精确因果建模的成本,并识别出问题可简化的条件。为验证理论,我们构建了一个反映真实世界因果结构的可控合成数据集,并进行了广泛实验。最后,我们开发了一个可无缝嵌入现有图学习流水线的因果建模增强模块,并通过全面对比实验展示了其有效性。

原文摘要 · Abstract (English)

Modeling causal relationships in graph representation learning remains a fundamental challenge. Existing approaches often draw on theories and methods from causal inference to identify causal subgraphs or mitigate confounders. However, due to the inherent complexity of graph-structured data, these approaches frequently aggregate diverse graph elements into single causal variables, an operation that risks violating the core assumptions of causal inference. In this work, we prove that such aggregation compromises causal validity. Building on this conclusion, we propose a theoretical model grounded in the smallest indivisible units of graph data to ensure that the causal validity is guaranteed. With this model, we further analyze the costs of achieving precise causal modeling in graph representation learning and identify the conditions under which the problem can be simplified. To empirically support our theory, we construct a controllable synthetic dataset that reflects realworld causal structures and conduct extensive experiments for validation. Finally, we develop a causal modeling enhancement module that can be seamlessly integrated into existing graph learning pipelines, and we demonstrate its effectiveness through comprehensive comparative experiments.

因果推断图神经网络可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。