arXiv:2504.14221cs.CV2025-04CVPR被引 19

构建首个融合真2D/伪3D/3D的工业缺陷检测数据集,提升多模态识别能力。

Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly Detection

  • 引入光度立体生成伪3D模态,融合RGB、点云与伪3D信息
  • 覆盖20类缺陷,包含微米级精度3D点云和高分辨率图像
  • 适用于工业视觉中多模态异常检测研究者

工业异常检测(IAD)的复杂性日益增加,多模态检测方法成为机器视觉研究热点。然而,专为IAD设计的多模态数据集仍十分有限。尽管MVTec 3D等开创性数据集已引入RGB+3D数据,但在真实工业环境中的规模与分辨率仍显不足。为此,我们提出Real-IAD D3,一个高精度多模态数据集,首次融合了通过光度立体法生成的伪3D模态,以及高分辨率RGB图像与微米级3D点云。该数据集涵盖20个类别,包含更细粒度缺陷、多样异常及更大规模,为多模态IAD提供挑战性基准。同时,我们提出一种有效融合方法,整合RGB、点云与伪3D深度信息,发挥各模态互补优势,显著提升检测鲁棒性与整体性能。实验验证了多模态信息对性能的关键作用。数据集与代码已公开:https://realiad4ad.github.io/Real-IAD D3

原文摘要 · Abstract (English)

The increasing complexity of industrial anomaly detection (IAD) has positioned multimodal detection methods as a focal area of machine vision research. However, dedicated multimodal datasets specifically tailored for IAD remain limited. Pioneering datasets like MVTec 3D have laid essential groundwork in multimodal IAD by incorporating RGB+3D data, but still face challenges in bridging the gap with real industrial environments due to limitations in scale and resolution. To address these challenges, we introduce Real-IAD D3, a high-precision multimodal dataset that uniquely incorporates an additional pseudo3D modality generated through photometric stereo, alongside high-resolution RGB images and micrometer-level 3D point clouds. Real-IAD D3 features finer defects, diverse anomalies, and greater scale across 20 categories, providing a challenging benchmark for multimodal IAD Additionally, we introduce an effective approach that integrates RGB, point cloud, and pseudo-3D depth information to leverage the complementary strengths of each modality, enhancing detection performance. Our experiments highlight the importance of these modalities in boosting detection robustness and overall IAD performance. The dataset and code are publicly accessible for research purposes at https://realiad4ad.github.io/Real-IAD D3

工业检测多模态伪3D点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。