arXiv:2504.09635cs.AIstat.ME2025-04被引 1

提出两阶段可解释匹配框架,提升因果推断中处理组与对照组的可比性。

A Two-Stage Interpretable Matching Framework for Causal Inference

  • 先精确匹配所有协变量,再逐步剔除影响最小的混杂因素
  • 显著降低因果效应估计偏差,提升处理组与对照组多维重叠度
  • 适合医疗数据分析等需透明解释的高维观测数据场景

基于观测数据的因果推断中的匹配方法旨在构建协变量分布相似的处理组与对照组,从而减少混杂并实现处理效应的无偏估计。该匹配样本近似于随机对照试验(RCT),提升因果推断质量。本文提出一种新型两阶段可解释匹配(TIM)框架,实现透明、可解释的协变量匹配。第一阶段对所有可用协变量进行精确匹配;对于第一阶段无法匹配的处理与对照单位,进入第二阶段:通过迭代剔除最不重要的混杂变量,并在剩余协变量上重新尝试精确匹配,同时学习被剔除协变量的距离度量,以量化其与处理单元在相应分层中的接近程度。利用高质量匹配样本估计条件平均处理效应(CATE)。为验证性能,我们在具有不同关联结构和相关性的合成数据集上进行实验,通过评估处理效应估计偏差及匹配前后处理组与对照组的多变量重叠度来衡量效果。此外,将TIM应用于美国疾控中心(CDC)的真实医疗数据,估计高胆固醇对糖尿病的因果效应。结果表明,TIM能有效改善CATE估计,增强多变量重叠性,并在高维数据中表现良好,是观测数据中稳健的因果推断工具。

原文摘要 · Abstract (English)

Matching in causal inference from observational data aims to construct treatment and control groups with similar distributions of covariates, thereby reducing confounding and ensuring an unbiased estimation of treatment effects. This matched sample closely mimics a randomized controlled trial (RCT), thus improving the quality of causal estimates. We introduce a novel Two-stage Interpretable Matching (TIM) framework for transparent and interpretable covariate matching. In the first stage, we perform exact matching across all available covariates. For treatment and control units without an exact match in the first stage, we proceed to the second stage. Here, we iteratively refine the matching process by removing the least significant confounder in each iteration and attempting exact matching on the remaining covariates. We learn a distance metric for the dropped covariates to quantify closeness to the treatment unit(s) within the corresponding strata. We used these high- quality matches to estimate the conditional average treatment effects (CATEs). To validate TIM, we conducted experiments on synthetic datasets with varying association structures and correlations. We assessed its performance by measuring bias in CATE estimation and evaluating multivariate overlap between treatment and control groups before and after matching. Additionally, we apply TIM to a real-world healthcare dataset from the Centers for Disease Control and Prevention (CDC) to estimate the causal effect of high cholesterol on diabetes. Our results demonstrate that TIM improves CATE estimates, increases multivariate overlap, and scales effectively to high-dimensional data, making it a robust tool for causal inference in observational data.

因果推断匹配方法可解释性医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。