通过多模态导师学习提升3D异常检测精度
Mentor3AD: Feature Reconstruction-based 3D Anomaly Detection via Multi-modality Mentor Learning
- 融合RGB与3D特征生成导师特征,指导重建
- 在MVTec 3D-AD和Eyecandies上达到新高
- 适合需要高精度3D异常检测的工业场景
多模态特征重建是3D异常检测的有前景方法,可利用双模态的互补信息。本文进一步提出多模态导师学习,通过融合中间特征以更清晰地区分正常与异常差异。为此,我们提出Mentor3AD方法,利用不同模态共享特征提取更有效表示,并指导特征重建,从而提升检测性能。具体包括:融合模块(MFM)将RGB与3D模态特征合并生成导师特征;引导模块(MGM)基于导师特征实现跨模态重建;投票模块(VM)更准确生成最终异常分数。在MVTec 3D-AD和Eyecandies数据集上的对比与消融实验验证了方法有效性。
原文摘要 · Abstract (English)
Multimodal feature reconstruction is a promising approach for 3D anomaly detection, leveraging the complementary information from dual modalities. We further advance this paradigm by utilizing multi-modal mentor learning, which fuses intermediate features to further distinguish normal from feature differences. To address these challenges, we propose a novel method called Mentor3AD, which utilizes multi-modal mentor learning. By leveraging the shared features of different modalities, Mentor3AD can extract more effective features and guide feature reconstruction, ultimately improving detection performance. Specifically, Mentor3AD includes a Mentor of Fusion Module (MFM) that merges features extracted from RGB and 3D modalities to create a mentor feature. Additionally, we have designed a Mentor of Guidance Module (MGM) to facilitate cross-modal reconstruction, supported by the mentor feature. Lastly, we introduce a Voting Module (VM) to more accurately generate the final anomaly score. Extensive comparative and ablation studies on MVTec 3D-AD and Eyecandies have verified the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。