为工业钢板缺陷检测提供带统计保证的可靠预测,避免误判漏判。
Conformal Segmentation in Industrial Surface Defect Detection with Statistical Guarantees
- 基于校准数据构建误差量化损失函数,确保预测结果可信赖。
- 在不同风险水平下,缺陷区域预测误差率严格控制在预设阈值内。
- 适合对可靠性要求高的工业质检场景,尤其适用于小样本或不均衡数据。
工业中钢板表面缺陷会显著影响其使用寿命并带来安全隐患。传统检测依赖人工,效率低且成本高;虽有基于卷积神经网络(如Mask R-CNN)的自动化方法,但因训练时标注不确定性和过拟合问题,新样本检测易出现偏差,可靠性不足。为此,本文通过满足独立同分布(i.i.d)条件的校准数据评估模型实际性能,为每个校准样本定义损失函数,量化检测误差率(如召回率补集与假阳性发现率)。据此推导出用户设定风险水平下的统计严格阈值,用于识别测试图像中高概率缺陷像素,构建预测集(如缺陷区域)。该方法确保测试集上期望误差率(均值误差率)严格低于预设风险水平。同时,观察到测试集上预测集平均大小与风险水平呈负相关,建立了衡量模型不确定性的统计指标。此外,该方法在不同校准-测试划分比例下仍能稳健高效控制预期误差率,验证了其适应性与实用性。
原文摘要 · Abstract (English)
In industrial settings, surface defects on steel can significantly compromise its service life and elevate potential safety risks. Traditional defect detection methods predominantly rely on manual inspection, which suffers from low efficiency and high costs. Although automated defect detection approaches based on Convolutional Neural Networks(e.g., Mask R-CNN) have advanced rapidly, their reliability remains challenged due to data annotation uncertainties during deep model training and overfitting issues. These limitations may lead to detection deviations when processing the given new test samples, rendering automated detection processes unreliable. To address this challenge, we first evaluate the detection model's practical performance through calibration data that satisfies the independent and identically distributed (i.i.d) condition with test data. Specifically, we define a loss function for each calibration sample to quantify detection error rates, such as the complement of recall rate and false discovery rate. Subsequently, we derive a statistically rigorous threshold based on a user-defined risk level to identify high-probability defective pixels in test images, thereby constructing prediction sets (e.g., defect regions). This methodology ensures that the expected error rate (mean error rate) on the test set remains strictly bounced by the predefined risk level. Additionally, we observe a negative correlation between the average prediction set size and the risk level on the test set, establishing a statistically rigorous metric for assessing detection model uncertainty. Furthermore, our study demonstrates robust and efficient control over the expected test set error rate across varying calibration-to-test partitioning ratios, validating the method's adaptability and operational effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。