提出动态控制图监测图像分类概率预测的校准状态,可早期发现模型性能退化。
Monitoring the calibration of probability forecasts with an application to concept drift detection involving image classification
- 基于累积和统计量设计动态阈值,持续监控概率预测校准情况。
- 无需访问模型内部即可检测校准偏差,适用于实时场景。
- 适合部署在工业、安防等需长期稳定预测的图像识别系统中。
图像分类的机器学习方法取得了显著进展,例如卷积神经网络在工业、国防等领域实现了高精度分类。然而,模型预测概率的校准性仍面临挑战:理想的校准是指预测概率与实际事件发生率一致。尽管已有多种评估和校正校准性的方法,但针对随时间推移持续监测模型校准状态的研究较少。本文提出一种基于累积和的动态限值方法,可用于传统过程监控及概念漂移检测。该方法能早期发现影响图像分类性能的运行环境变化。所提图表广泛适用于需要长期监控概率预测校准性的场景,且仅需概率输出与实际结果,无需访问模型内部结构。
原文摘要 · Abstract (English)
Machine learning approaches for image classification have led to impressive advances in that field. For example, convolutional neural networks are able to achieve remarkable image classification accuracy across a wide range of applications in industry, defense, and other areas. While these machine learning models boast impressive accuracy, a related concern is how to assess and maintain calibration in the predictions these models make. A classification model is said to be well calibrated if its predicted probabilities correspond with the rates events actually occur. While there are many available methods to assess machine learning calibration and recalibrate faulty predictions, less effort has been spent on developing approaches that continually monitor predictive models for potential loss of calibration as time passes. We propose a cumulative sum-based approach with dynamic limits that enable detection of miscalibration in both traditional process monitoring and concept drift applications. This enables early detection of operational context changes that impact image classification performance in the field. The proposed chart can be used broadly in any situation where the user needs to monitor probability predictions over time for potential lapses in calibration. Importantly, our method operates on probability predictions and event outcomes and does not require under-the-hood access to the machine learning model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。