arXiv:2412.03913cs.LGcs.IR2024-12中稿 · WSDM 2025被引 7

提出GDC模型,分离网络数据中的调整变量与混杂因素,提升因果推断准确率

Graph Disentangle Causal Model: Enhancing Causal Inference in Networked Observational Data

  • 通过因果解耦模块将特征分为调整变量和混杂因素
  • 在两个网络数据集上显著降低混杂偏倚,提升个体处理效应估计精度
  • 适合处理特征分布差异大、存在隐藏混杂的复杂网络数据

从观测数据中估计个体处理效应(ITE)是多个领域的重要任务。然而,现有方法常忽略个体层面未观测到的隐藏混杂因素。研究者利用图神经网络聚合邻居特征以捕捉隐藏混杂因素,并通过最小化处理组与对照组混杂表示的差异来缓解混杂偏倚。尽管取得成功,但在实际场景中,往往将所有特征视为混杂因素,且处理组与对照组间特征分布差异显著。将调整变量误认为混杂因素并强制混杂表示严格平衡,可能损害结果预测效果。为此,我们提出新型框架——图解耦因果模型(GDC),用于网络环境下的ITE估计。GDC采用因果解耦模块,将单元特征分离为调整变量与混杂因素表示;设计包含三个不同图聚合器的图聚合模块,获取调整变量、混杂因素及反事实混杂因素表示;最后使用因果约束模块确保解耦表示为真实因果因子。在两个网络数据集上的全面实验验证了所提方法的有效性。

原文摘要 · Abstract (English)

Estimating individual treatment effects (ITE) from observational data is a critical task across various domains. However, many existing works on ITE estimation overlook the influence of hidden confounders, which remain unobserved at the individual unit level. To address this limitation, researchers have utilized graph neural networks to aggregate neighbors' features to capture the hidden confounders and mitigate confounding bias by minimizing the discrepancy of confounder representations between the treated and control groups. Despite the success of these approaches, practical scenarios often treat all features as confounders and involve substantial differences in feature distributions between the treated and control groups. Confusing the adjustment and confounder and enforcing strict balance on the confounder representations could potentially undermine the effectiveness of outcome prediction. To mitigate this issue, we propose a novel framework called the \textit{Graph Disentangle Causal model} (GDC) to conduct ITE estimation in the network setting. GDC utilizes a causal disentangle module to separate unit features into adjustment and confounder representations. Then we design a graph aggregation module consisting of three distinct graph aggregators to obtain adjustment, confounder, and counterfactual confounder representations. Finally, a causal constraint module is employed to enforce the disentangled representations as true causal factors. The effectiveness of our proposed method is demonstrated by conducting comprehensive experiments on two networked datasets.

因果推断图神经网络混杂因素ITE估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。