通过因果图拓扑发现同质子群体,纠正严重混杂偏差。
CaSPECT: Discovering Causally Homogeneous Subgroups via Directed Spectral Clustering

- 基于有向图拓扑定义相似性,不依赖协变量空间聚类。
- 在多个真实数据集上识别出显著正向处理效应,效果优于传统方法。
- 无需预设倾向得分模型,适合因果推断中复杂混杂场景。
我们提出 CaSPECT,一种基于观测数据发现因果同质子群体的因果谱聚类框架。不同于在协变量空间中聚类,CaSPECT 通过学习的有向无环图(DAG)拓扑定义相似性;采用自助稳定化的 PC 算法恢复因果骨架;提出新颖的定向验证评分(OVS),结合 PC 自助证据与 DirectLiNGAM 实现稳健边向定向;有向边权重由后门可识别的平均处理效应确定,采用 OLS 或双重机器学习估计。Chung 的有向拉普拉斯矩阵提供谱嵌入,使彼此接近的个体共享相同的因果传播路径。我们建立了整个流程的几乎必然一致性,并通过受控模拟研究以及 LaLonde CPS1、IHDP 和 401(k) 数据集验证了该方法,在其中 CaSPECT 在因果可比子群体中识别出正向且统计显著的处理效应,有效纠正了严重混杂,且无需预先指定倾向得分模型。
原文摘要 · Abstract (English)
We propose \textbf{CaSPECT}, a causal spectral clustering framework for discovering causally homogeneous subgroups from observational data. Rather than clustering in covariate space, CaSPECT defines similarity through the topology of a learned directed acyclic graph (DAG); a bootstrap-stabilised PC algorithm recovers the causal skeleton; a novel \emph{Orientation Validation Score} (OVS) combines PC bootstrap evidence with DirectLiNGAM to orient edges robustly; directed edges are weighted by backdoor-identified average treatment effects estimated via OLS or double machine learning. Chung's directed Laplacian provides a spectral embedding in which individuals close together share the same causal propagation pathways. We establish almost-sure consistency of the full pipeline and validate the method through a controlled simulation study and on LaLonde CPS1, IHDP, and 401(k) datasets, where CaSPECT recovers a positive and statistically significant treatment effect within the causally comparable subpopulation and corrects for severe confounding without requiring a pre-specified propensity score model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。