arXiv:2412.12910stat.MLcs.LG2024-12NeurIPS被引 12

无标签环境下检测模型性能下降的分布偏移

Sequential Harmful Shift Detection Without Labels

  • 用误差估计器预测值作为真实误差的代理
  • 在多种分布偏移下检测效能高且误报率低
  • 适合无标签数据的持续部署场景

我们提出一种新方法,用于在连续生产环境中检测对机器学习模型性能产生负面影响的分布偏移,无需访问真实标签数据。该方法基于Podkopaev和Ramdas [2022]的工作,后者在有标签的情况下追踪模型错误随时间的变化。我们的方案通过训练一个误差估计器,利用其预测结果作为真实误差的代理,从而在无标签条件下实现检测。实验表明,该方法在各种分布偏移(包括协变量偏移、标签偏移以及地理和时间上的自然偏移)下均具备高检测力和良好的误报控制能力。

原文摘要 · Abstract (English)

We introduce a novel approach for detecting distribution shifts that negatively impact the performance of machine learning models in continuous production environments, which requires no access to ground truth data labels. It builds upon the work of Podkopaev and Ramdas [2022], who address scenarios where labels are available for tracking model errors over time. Our solution extends this framework to work in the absence of labels, by employing a proxy for the true error. This proxy is derived using the predictions of a trained error estimator. Experiments show that our method has high power and false alarm control under various distribution shifts, including covariate and label shifts and natural shifts over geography and time.

分布偏移无监督检测模型监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。