arXiv:2506.05047cs.LG2025-06NeurIPS被引 3

不依赖标签,实时检测模型部署后性能下降

Reliably Detecting Model Failures in Deployment Without Labels

  • 基于模型预测分歧度设计监控算法
  • 非退化数据漂移下误报率低,退化时真阳性率高
  • 适合医疗等高风险场景的模型健康监测

数据分布随时间变化,动态环境中运行的模型需定期重训练。但缺乏标签时,判断何时重训仍具挑战,因并非所有分布漂移都会降低模型性能。本文形式化并解决部署后性能劣化(PDD)监测问题。提出D3M算法,基于模型预测分歧设计,兼具高效性与实用性,在非退化漂移下保持低误报率,并在退化漂移下提供高真阳性率的样本复杂度保证。在标准基准和一个大规模真实世界内科数据集上的实验证明该框架有效,具备作为高风险机器学习流水线预警机制的可行性。

原文摘要 · Abstract (English)

The distribution of data changes over time; models operating in dynamic environments need retraining. But knowing when to retrain, without access to labels, is an open challenge since some, but not all shifts degrade model performance. This paper formalizes and addresses the problem of post-deployment deterioration (PDD) monitoring. We propose D3M, a practical and efficient monitoring algorithm based on the disagreement of predictive models, achieving low false positive rates under non-deteriorating shifts and provides sample complexity bounds for high true positive rates under deteriorating shifts. Empirical results on both standard benchmark and a real-world large-scale internal medicine dataset demonstrate the effectiveness of the framework and highlight its viability as an alert mechanism for high-stakes machine learning pipelines.

模型监控无监督检测部署优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。