改进倾向得分估计的校准策略,提升因果推断在小样本等困难场景下的稳定性与准确性。
Calibration Strategies for Robust Causal Estimation: Theoretical and Empirical Insights on Propensity Score-Based Estimators
- 采用分样本校准框架,优化倾向得分估计的稳健性。
- 校准显著降低IPW estimator方差并缓解偏差,尤其在小样本中效果明显。
- 对复杂模型(如梯度提升)有稳定作用,适合高维或数据不平衡场景研究者。
基于倾向得分的估计方法(如逆概率加权IPW和双重机器学习DML)的性能高度依赖于数据划分策略。本文扩展了倾向得分估计的校准技术,在重叠不足、样本量小或数据不平衡等挑战性场景下提升了倾向得分的鲁棒性。首先,我们对DML框架中校准估计器的性质进行了理论分析,重点探讨了分样本方案对有效因果推断的影响。其次,通过大量模拟实验发现,校准能有效降低基于逆概率的估计器方差,并缓解IPW的偏差,即使在小样本条件下亦然。特别地,校准增强了灵活学习器(如梯度提升)的稳定性,同时保持DML的双重稳健性。关键发现是:即便不校准时方法表现良好,只要选择合适的分样本策略,加入校准步骤也不会损害性能。
原文摘要 · Abstract (English)
The partitioning of data for estimation and calibration critically impacts the performance of propensity score based estimators like inverse probability weighting (IPW) and double/debiased machine learning (DML) frameworks. We extend recent advances in calibration techniques for propensity score estimation, improving the robustness of propensity scores in challenging settings such as limited overlap, small sample sizes, or unbalanced data. Our contributions are twofold: First, we provide a theoretical analysis of the properties of calibrated estimators in the context of DML. To this end, we refine existing calibration frameworks for propensity score models, with a particular emphasis on the role of sample-splitting schemes in ensuring valid causal inference. Second, through extensive simulations, we show that calibration reduces variance of inverse-based propensity score estimators while also mitigating bias in IPW, even in small-sample regimes. Notably, calibration improves stability for flexible learners (e.g., gradient boosting) while preserving the doubly robust properties of DML. A key insight is that, even when methods perform well without calibration, incorporating a calibration step does not degrade performance, provided that an appropriate sample-splitting approach is chosen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。