arXiv:2606.14506stat.MLcs.LG2026-06被引 1

评估模型在分布偏移和选择性标签下的表现,提升部署前预测可靠性。

Beyond the Training Distribution: Evaluating Predictions Under Distribution Shift and Selection Bias

  • 提出双机器学习方法,估计任意黑箱模型在新环境中的风险。
  • 实验显示该方法比单独处理任一问题的基准更准确追踪真实风险。
  • 适合医疗等高风险场景中模型上线前的风险评估使用。

在算法影响决策时,预判模型在新环境中的表现对防止危害至关重要。常见性能下降原因包括:(i) 预测变量分布偏移(协变量偏移),即目标分布与源分布不同;(ii) 选择性标注,即结果可观测性依赖于历史决策。本文研究在协变量偏移与基于观测特征的选择性标注共同存在下的部署前模型评估。提出一种双机器学习程序,用于在一般损失函数下估计任意黑箱模型的目标风险。在标准假设下证明该估计量的可识别性,并基于目标风险的影响函数推导出偏差校正估计器。最终在eICU电子健康记录数据库上进行实验,结果表明,该方法比仅处理选择性标签或协变量偏移的单一方法,以及结合传统插值法的基线方法,更能准确追踪真实目标风险。

原文摘要 · Abstract (English)

Understanding how a prediction model will perform in a new environment before deployment is essential to preventing harm when algorithms inform decision-making. Two common sources of model performance degradation are (i) covariate shift, where the target covariate distribution differs from the source, and (ii) selective labels, where the observability of outcomes depends on historical decisions. We study pre-deployment model evaluation under the joint presence of covariate shift and labeling of outcomes selectively based on observed features. In particular, we present a double machine learning procedure for estimating the target risk of an arbitrary black-box prediction model under a general loss function. We show identification of this estimand under standard assumptions and derive a bias-corrected estimator based on the influence function of the target risk. Finally, we evaluate our estimator through experiments using the eICU electronic health records database, showing that it tracks the true target risk more accurately than methods that address either selective labels or covariate shift alone, as well as baselines that combine standard plug-in approaches.

模型评估分布偏移选择性标注医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。