统一多类别工业缺陷检测,用参考特征动态过滤噪声。
Uni-RCM: Unified Reference-guided Cross-modal Mapping for Multi-Class Anomaly Detection

- 引入可学习参考特征,动态过滤类别特有噪声。
- 在MVTec-3D AD上实现像素级定位的顶尖性能。
- 适合需要统一模型处理多种产品的场景。
多模态工业异常检测通常为每个产品类别单独建模,严重限制了实际可扩展性。转向统一处理多种类别的范式时,因类别间干扰和特征流形混淆,检测精度常下降。为此,本文提出统一参考引导跨模态映射框架Uni-RCM。核心是参考引导模块,通过引入可学习参考特征,动态过滤类别特异性噪声,捕捉不同模态间的共性。此外,采用离线残差量化器,通过多级码本刻画正常分布。在MVTec-3D AD数据集上的大量实验表明,该方法在多类别设定下达到当前最优的图像级检测与像素级定位性能。
原文摘要 · Abstract (English)
Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. When shifting to a unified paradigm that handles diverse classes simultaneously, detection accuracy often degrades due to inter-class interference and feature manifold confusion. To overcome these challenges, we propose a Unified Reference guided Cross-modal Mapping framework, named Uni-RCM. At its core, we propose a reference guide block to dynamically filter out category-specific noise by introducing a learnable reference feature, which captures the commonalities across different modalities. Besides, an offline residual quantizer is proposed to characterize the normal distribution by multiple cascaded codebooks. Extensive evaluations on the MVTec-3D AD dataset demonstrate the state-of-the-art performance in the challenging multi-class setting and in terms of image-level detection and pixel-level localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。