用跨模态正则约束提升无监督多类异常检测精度
CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly Detection
- 引入与类别无关的可学习提示,统一正常视觉模式的语义表征
- 在MVTec AD和VisA上达到领先性能,异常定位更准确
- 适合需要高鲁棒性异常检测的工业质检场景
现有基于蒸馏的无监督方法依赖编码与解码特征差异定位异常区域,但仅在正常样本上训练的解码器仍能较好重构异常块特征,导致性能下降。这一问题在多类异常检测中尤为严重。我们将其归因于解码器的过度泛化(OG):多类训练中块模式多样性增加,虽提升了对正常块的泛化能力,也无意中扩大了对异常块的泛化范围。为此,我们提出一种新方法,利用与类别无关的可学习提示捕捉各类视觉模式下的共性正常语义,并引导解码特征向正常文本表征靠拢,抑制解码器对异常模式的过度泛化。为进一步提升性能,还引入门控专家混合模块,使模型能更好处理多样块模式,减少多类训练中的相互干扰。实验表明,该方法在MVTec AD和VisA数据集上表现优异,验证了其有效性。
原文摘要 · Abstract (English)
Existing unsupervised distillation-based methods rely on the differences between encoded and decoded features to locate abnormal regions in test images. However, the decoder trained only on normal samples still reconstructs abnormal patch features well, degrading performance. This issue is particularly pronounced in unsupervised multi-class anomaly detection tasks. We attribute this behavior to over-generalization(OG) of decoder: the significantly increasing diversity of patch patterns in multi-class training enhances the model generalization on normal patches, but also inadvertently broadens its generalization to abnormal patches. To mitigate OG, we propose a novel approach that leverages class-agnostic learnable prompts to capture common textual normality across various visual patterns, and then apply them to guide the decoded features towards a normal textual representation, suppressing over-generalization of the decoder on abnormal patterns. To further improve performance, we also introduce a gated mixture-of-experts module to specialize in handling diverse patch patterns and reduce mutual interference between them in multi-class training. Our method achieves competitive performance on the MVTec AD and VisA datasets, demonstrating its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。