arXiv:2508.05215cs.LGstat.ME2025-08中稿 · Frontiers in Appli…

提出新权重方法DFW,让观察性研究更接近随机实验。

DFW: A Novel Weighting Scheme for Covariate Balancing and Treatment Effect Estimation

  • 用去混杂因子构建稳定权重,避免传统方法的极端值问题。
  • 在真实与合成数据上,显著提升协变量平衡与因果效应估计精度。
  • 适合处理有混杂因素的医疗、政策评估等现实场景分析。

从观察性数据中估计因果效应面临选择偏差挑战,导致处理组间协变量分布不平衡。基于倾向得分的加权方法常用于模拟随机对照试验(RCT),但其效果高度依赖于观测数据和倾向得分估计的准确性。例如,逆倾向得分加权(IPW)根据倾向得分的倒数分配权重,在倾向得分方差较大时(源于数据或模型误设)会产生不稳定的权重,从而削弱对选择偏差的处理能力并影响因果效应估计。为此,我们提出去混杂因子加权(DFW),一种新型基于倾向得分的方法,利用去混杂因子构造稳定且有效的样本权重。DFW优先考虑混杂程度较低的样本,同时降低高度混杂样本的影响,生成更接近RCT的伪总体。该方法保证权重有界、方差更低,并改善协变量平衡。尽管DFW针对二元处理设计,但可自然扩展至多处理场景,因去混杂因子基于每个样本实际接受治疗的概率计算。在真实世界基准和合成数据集上的大量实验表明,DFW在协变量平衡与因果效应估计方面优于现有方法,包括IPW和CBPS。

原文摘要 · Abstract (English)

Estimating causal effects from observational data is challenging due to selection bias, which leads to imbalanced covariate distributions across treatment groups. Propensity score-based weighting methods are widely used to address this issue by reweighting samples to simulate a randomized controlled trial (RCT). However, the effectiveness of these methods heavily depends on the observed data and the accuracy of the propensity score estimator. For example, inverse propensity weighting (IPW) assigns weights based on the inverse of the propensity score, which can lead to instable weights when propensity scores have high variance-either due to data or model misspecification-ultimately degrading the ability of handling selection bias and treatment effect estimation. To overcome these limitations, we propose Deconfounding Factor Weighting (DFW), a novel propensity score-based approach that leverages the deconfounding factor-to construct stable and effective sample weights. DFW prioritizes less confounded samples while mitigating the influence of highly confounded ones, producing a pseudopopulation that better approximates a RCT. Our approach ensures bounded weights, lower variance, and improved covariate balance.While DFW is formulated for binary treatments, it naturally extends to multi-treatment settings, as the deconfounding factor is computed based on the estimated probability of the treatment actually received by each sample. Through extensive experiments on real-world benchmark and synthetic datasets, we demonstrate that DFW outperforms existing methods, including IPW and CBPS, in both covariate balancing and treatment effect estimation.

因果推断倾向得分加权方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。