arXiv:2602.08892stat.MLcs.LG2026-02被引 1

模型评估常高估政策效果,真实收益可能为零。

Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching

  • 用合成数据模拟难民匹配,检验模型评估方法的可靠性
  • 即使条件理想,模型仍报告60%虚假提升,真实为零
  • 警示决策者警惕模型评估的乐观偏差,尤其在政策制定中

数据驱动决策的核心挑战在于准确评估政策效果——确保学习到的决策策略能实现承诺的收益。目前主流方法是基于模型的评估,即从数据中估计模型以推断反事实结果。但该方法因‘赢家诅咒’导致对真实收益的过度乐观估计。我们调研了近十年《管理科学》上55篇相关论文,发现除两篇外均采用此有缺陷的方法。常见辩护包括:模型准确稳定、历史数据随机分配、模型设定恰当、使用样本分割。然而我们证明,这些条件组合无法避免赢家诅咒。理论分析显示,即使所有条件满足,仍可能出现大幅虚假收益。仿真研究基于真实的难民匹配问题构建合成环境(与现实设置高度匹配),但设计为任何分配策略都无法提高预期就业率。模型评估方法仍报告约60%的稳定增益,与文献中22%-75%的报告值相当,而真实效应为零。结果强烈质疑基于模型的评估方法的有效性。

原文摘要 · Abstract (English)

A major challenge in data-driven decision-making is accurate policy evaluation-i.e., guaranteeing that a learned decision-making policy achieves the promised benefits. A popular strategy is model-based policy evaluation, which estimates a model from data to infer counterfactual outcomes. This strategy is known to produce unwarrantedly optimistic estimates of the true benefit due to the winner's curse. We searched the recent literature on data-driven decision-making, identifying a sample of 55 papers published in the Management Science in the past decade; all but two relied on this flawed methodology. Several common justifications are provided: (1) the estimated models are accurate, stable, and well-calibrated, (2) the historical data uses random treatment assignment, (3) the model family is well-specified, and (4) the evaluation methodology uses sample splitting. Unfortunately, we show that no combination of these justifications avoids the winner's curse. First, we provide a theoretical analysis demonstrating that the winner's curse can cause large, spurious reported benefits even when all these justifications hold. Second, we perform a simulation study based on the recent and consequential data-driven refugee matching problem. We construct a synthetic refugee matching environment (calibrated to closely match the real setting) but designed so that no assignment policy can improve expected employment compared to random assignment. Model-based methods report large, stable gains of around 60% even when the true effect is zero; these gains are on par with improvements of 22-75% reported in the literature. Our results provide strong evidence against model-based evaluation.

政策评估赢家诅咒虚假提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。