解决工业表面缺陷检测中传感器缺失问题,提升多模态融合鲁棒性。
Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability

- 设计跨模态提示机制,应对模态缺失时的信息不一致与模式变换。
- 在0.7总缺失率下,实现73.83% I-AUROC和93.05% P-AUROC,优于现有方法。
- 适合传感器不稳定场景的工业缺陷检测系统部署。
多模态工业表面缺陷检测(MISDD)通过融合RGB与3D模态识别工业产品缺陷。本文聚焦于传感器不确定性导致的模态缺失问题。为此,提出跨模态提示学习:包括跨模态一致性提示以维持双视觉模态信息一致性;模态特定提示以适应不同输入模式;缺失感知提示以补偿动态模态缺失带来的信息空缺。此外,提出对称对比学习,利用文本模态作为双视觉模态融合桥梁。设计一对对立文本提示生成二元语义,通过三模态对比预训练完成多模态学习。实验表明,所提方法在RGB与3D模态总缺失率为0.7时,取得73.83% I-AUROC与93.05% P-AUROC,分别超越当前最优方法3.84%与5.58%,且在不同缺失类型与率下均表现更优。源代码将发布于https://github.com/SvyJ/MISDD-MM。
原文摘要 · Abstract (English)
Multimodal industrial surface defect detection (MISDD) aims to identify and locate defect in industrial products by fusing RGB and 3D modalities. This article focuses on modality-missing problems caused by uncertain sensors availability in MISDD. In this context, the fusion of multiple modalities encounters several troubles, including learning mode transformation and information vacancy. To this end, we first propose cross-modal prompt learning, which includes: i) the cross-modal consistency prompt serves the establishment of information consistency of dual visual modalities; ii) the modality-specific prompt is inserted to adapt different input patterns; iii) the missing-aware prompt is attached to compensate for the information vacancy caused by dynamic modalities-missing. In addition, we propose symmetric contrastive learning, which utilizes text modality as a bridge for fusion of dual vision modalities. Specifically, a paired antithetical text prompt is designed to generate binary text semantics, and triple-modal contrastive pre-training is offered to accomplish multimodal learning. Experiment results show that our proposed method achieves 73.83% I-AUROC and 93.05% P-AUROC with a total missing rate 0.7 for RGB and 3D modalities (exceeding state-of-the-art methods 3.84% and 5.58% respectively), and outperforms existing approaches to varying degrees under different missing types and rates. The source code will be available at https://github.com/SvyJ/MISDD-MM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。