用教师编码器+学生解码器去噪,提升异常检测分割精度
Teacher Encoder-Student Decoder Denoising Guided Segmentation Network for Anomaly Detection
- 教师编码器指导学生解码器,结合多尺度特征融合增强学习
- 在MVTec AD数据集上达到98.9%图像级AUC和76.4%像素级精确率
- 无需标注,自动生成异常掩码,适合工业缺陷检测场景
视觉异常检测是一项极具挑战性的任务,通常被归类为一类分类与分割问题。近期研究显示,学生-教师(S-T)框架能有效应对该挑战。然而,大多数S-T框架仅依赖预训练教师网络指导学生网络学习多尺度相似特征,忽视了学生网络通过多尺度特征融合提升学习能力的潜力。本研究提出一种新模型PFADSeg,将预训练教师网络、具有多尺度特征融合的去噪学生网络以及引导式异常分割网络整合为统一框架。通过采用独特的教师编码器-学生解码器去噪模式,模型增强了学生网络从教师特征中学习的能力。此外,引入自适应特征融合机制,训练出自监督分割网络,可自主合成异常掩码,显著提升检测性能。在广泛使用的MVTec AD数据集上的严格评估表明,PFADSeg表现优异,图像级AUC达98.9%,像素级平均精确率为76.4%,实例级平均精确率为78.7%。
原文摘要 · Abstract (English)
Visual anomaly detection is a highly challenging task, often categorized as a one-class classification and segmentation problem. Recent studies have demonstrated that the student-teacher (S-T) framework effectively addresses this challenge. However, most S-T frameworks rely solely on pre-trained teacher networks to guide student networks in learning multi-scale similar features, overlooking the potential of the student networks to enhance learning through multi-scale feature fusion. In this study, we propose a novel model named PFADSeg, which integrates a pre-trained teacher network, a denoising student network with multi-scale feature fusion, and a guided anomaly segmentation network into a unified framework. By adopting a unique teacher-encoder and student-decoder denoising mode, the model improves the student network's ability to learn from teacher network features. Furthermore, an adaptive feature fusion mechanism is introduced to train a self-supervised segmentation network that synthesizes anomaly masks autonomously, significantly increasing detection performance. Rigorous evaluations on the widely-used MVTec AD dataset demonstrate that PFADSeg exhibits excellent performance, achieving an image-level AUC of 98.9%, a pixel-level mean precision of 76.4%, and an instance-level mean precision of 78.7%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。