解决治疗数据缺失下的最优决策问题,提升政策制定可靠性
Learning Who to Treat When Treatment is Missing
- 提出基于缺失机制的新型估计方法,处理随机缺失的治疗数据
- 在假设正确时,性能接近理想情况,且样本量越大越精准
- 适合医疗、公共政策等治疗数据常缺失的实际场景
政策学习方法被广泛用于预算受限下的治疗分配。现有方法多假设治疗数据完整,但实际应用中常存在缺失,导致估计偏差和次优策略。本文将高效平均治疗效应(ATE)估计器扩展至政策价值与条件平均治疗效应(CATE)估计,在缺失随机(MAR)与完全条件随机缺失(MCCAR)设定下均有效。渐近效率分析表明:当MCCAR假设成立时,利用部分观测单位的MAR估计器既有效又比MCCAR估计器更高效。这为两种缺失情形下优先采用MAR估计提供了理论依据。综合实验使用合成与半合成数据集验证:正确指定缺失机制至关重要——误设估计器始终有偏,无论样本大小;而本文方法在假设满足时可达到近似最优性能。本工作为缺失治疗数据下的稳健政策学习提供了理论严谨且实证有效的工具。
原文摘要 · Abstract (English)
Policy learning methods are increasingly used to inform treatment allocation under budget constraints. Most proposed methods assume complete treatment data, yet applications frequently suffer from missingness that can bias estimates and lead to suboptimal policies. We address this gap by extending efficient estimators for average treatment effect (ATE) estimation to policy value and conditional average treatment effect (CATE) estimation under missing at random (MAR) and missing completely conditionally at random (MCCAR) treatment data. Through asymptotic efficiency analysis, we prove that the MAR estimator, which leverages partially-observed units, is both valid and more efficient than the MCCAR estimator when MCCAR assumptions hold. This result provides formal justification for preferring MAR-based estimation in policy learning under both missing data settings. Our comprehensive experiments using synthetic and semi-synthetic datasets confirm that correctly specifying the missingness mechanism is crucial: misspecified estimators remain biased regardless of sample size, while our estimators achieve near-oracle performance when assumptions are satisfied. Our work provides practitioners with theoretically grounded, empirically validated tools for robust policy learning in the presence of missing treatment data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。