用真实实验数据对比干预策略,发现基于效果预测的选人方式更优。
Comparing Targeting Strategies for Maximizing Social Welfare with Limited Resources
- 用真实随机试验数据比较不同选人方法,评估其社会福利提升效果。
- 当能较准确估计个体干预效果时,按效果预测选人比按风险选人效率高得多。
- 现有数据常不足以准确估计差异性干预效果,需改进数据收集与建模。
机器学习在人力资源、教育、发展等领域的有限资源干预中应用日益广泛。然而,模型应预测什么目标尚不明确:政策制定者通常缺乏随机对照试验(RCT)数据来准确估算谁更可能从干预中受益,而观察性数据则存在严重偏倚风险。实践中普遍采用‘风险导向选人’策略,即仅预测个体当前状态(非因果任务),并优先向高风险者提供帮助。但目前几乎无实证研究可指导社会领域中机器学习辅助干预策略的选择。本文利用5个真实世界的跨领域随机试验数据,实证评估了不同策略的效果。结果表明,当干预效果可被高精度估计时(通过提前部分观测结果模拟实现),基于干预效果的选人策略显著优于风险导向策略,即使效果估计存在偏差也成立。该结论在政策制定者偏好帮扶高风险人群的情境下依然有效。然而,在我们考察的多数实际RCT中,可用特征与数据仍不足以实现对异质性干预效果的准确估计。研究提示,基于效果的选人具有巨大潜力,但要实现这一潜力,必须在数据收集和模型训练上超越当前常见做法。
原文摘要 · Abstract (English)
Machine learning is increasingly used to select which individuals receive limited-resource interventions in domains such as human services, education, development, and more. However, it is often not apparent what the right quantity is for models to predict. Policymakers rarely have access to data from a randomized controlled trial (RCT) that would enable accurate estimates of which individuals would benefit more from the intervention, while observational data creates a substantial risk of bias in treatment effect estimates. Practitioners instead commonly use a technique termed ``risk-based targeting" where the model is just used to predict each individual's status quo outcome (an easier, non-causal task). Those with higher predicted risk are offered treatment. There is currently almost no empirical evidence to inform which choices lead to the most effective machine learning-informed targeting strategies in social domains. In this work, we use data from 5 real-world RCTs in a variety of domains to empirically assess such choices. We find that when treatment effects can be estimated with high accuracy (which we simulate by allowing the model to partially observe outcomes in advance), treatment effect based targeting substantially outperforms risk-based targeting, even when treatment effect estimates are biased. Moreover, these results hold even when the policymaker has strong normative preferences for assisting higher-risk individuals. However, the features and data actually available in most RCTs we examine do not suffice for accurate estimates of heterogeneous treatment effects. Our results suggest treatment effect targeting has significant potential benefits, but realizing these benefits requires improvements to data collection and model training beyond what is currently common in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。