解决因果推断中因历史数据不平衡导致的估计偏差问题
A Distributionally Robust Framework for Nuisance in Causal Effect Estimation
- 用对抗损失同时优化倾向得分不确定性和权重不稳定性
- 在合成与真实数据集上均优于现有方法,提升因果效应估计精度
- 适合处理存在历史选择偏误的医疗、政策评估等场景
因果推断需在处理组与对照组平衡分布下评估模型,但训练数据常因历史决策策略而失衡。传统方法多采用逆概率加权(IPW),需先估计倾向得分,面临两个关键挑战:倾向得分估计不准和极端权重带来的不稳定性。本文通过分解泛化误差,分离出倾向得分模糊性与统计不稳定性,并提出一种对抗性损失函数予以解决。方法结合分布鲁棒优化以应对倾向得分不确定性,以及基于加权Rademacher复杂度的权重正则化。在合成数据与真实世界数据集上的实验表明,该方法在多种评估指标上均持续优于现有方法。
原文摘要 · Abstract (English)
Causal inference requires evaluating models on balanced distributions between treatment and control groups, while training data often exhibits imbalance due to historical decision-making policies. Most conventional statistical methods address this distribution shift through inverse probability weighting (IPW), which requires estimating propensity scores as an intermediate step. These methods face two key challenges: inaccurate propensity estimation and instability from extreme weights. We decompose the generalization error to isolate these issues--propensity ambiguity and statistical instability--and address them through an adversarial loss function. Our approach combines distributionally robust optimization for handling propensity uncertainty with weight regularization based on weighted Rademacher complexity. Experiments on synthetic and real-world datasets demonstrate consistent improvements over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。