arXiv:2505.15215stat.MLcs.LG2025-05被引 2

通过变量剪枝与聚类,简化复杂因果数据融合模型。

Clustering and Pruning in Causal Data Fusion

  • 用剪枝和聚类预处理因果数据融合,减少变量冗余。
  • 推导多源数据下变量合并与删除的可识别性条件。
  • 适用于高维因果建模,尤其适合流行病学与社会科学。

数据融合是结合观测数据与实验数据以识别原本不可识别的因果效应的过程。尽管特定场景已有识别算法,但当变量在部分数据源中存在而其他源缺失时,do-演算仍是唯一通用工具。然而,随着变量增多和因果图复杂度上升,基于do-演算的方法面临计算挑战。为此,本文提出将剪枝(移除不必要变量)与聚类(合并变量)作为因果数据融合的预处理操作。我们推广了单源数据的早期结果,推导出多源数据下剪枝与聚类的适用条件,并给出基于较小图推断较大图中因果效应可识别性或不可识别性的充分条件,同时说明如何获得可识别因果效应的识别函数。流行病学与社会科学研究中的实例验证了方法的有效性。

原文摘要 · Abstract (English)

Data fusion, the process of combining observational and experimental data, can enable the identification of causal effects that would otherwise remain non-identifiable. Although identification algorithms have been developed for specific scenarios, do-calculus remains the only general-purpose tool for causal data fusion, particularly when variables are present in some data sources but not others. However, approaches based on do-calculus may encounter computational challenges as the number of variables increases and the causal graph grows in complexity. Consequently, there exists a need to reduce the size of such models while preserving the essential features. For this purpose, we propose pruning (removing unnecessary variables) and clustering (combining variables) as preprocessing operations for causal data fusion. We generalize earlier results on a single data source and derive conditions for applying pruning and clustering in the case of multiple data sources. We give sufficient conditions for inferring the identifiability or non-identifiability of a causal effect in a larger graph based on a smaller graph and show how to obtain the corresponding identifying functional for identifiable causal effects. Examples from epidemiology and social science demonstrate the use of the results.

因果推断数据融合变量聚类剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。