arXiv:2503.04997cs.CV2025-03被引 6

构建首个大规模工业缺陷检测数据集,融合真实与合成缺陷。

ISP-AD: A Large-Scale Real-World Dataset for Advancing Industrial Anomaly Detection with Synthetic and Real Defects

论文配图:ISP-AD: A Large-Scale Real-World Dataset for Advancing Industrial Anomaly Detection with Synthetic and Real Defects
图 1 · 摘自论文原文
  • 构建含真实与合成缺陷的工业级数据集,覆盖复杂成像条件。
  • 少量真实缺陷可显著提升模型对未知缺陷的泛化能力。
  • 适合研究工业无监督/自监督/有监督异常检测的学者使用。

机器学习驱动的自动视觉检测在实现工业零缺陷目标中至关重要。现有异常检测研究受限于难以获取能反映复杂缺陷形态及非理想成像条件的数据集,而多数公开数据集偏向理想成像条件,高估了实际适用性。为此,我们提出工业屏幕印刷异常检测数据集(ISP-AD),包含嵌入在结构化图案中的微小、弱对比度缺陷,且设计变化范围大。据我们所知,它是目前最大的公开工业数据集,包含工厂实采的真实缺陷和合成缺陷。实验验证了混合训练策略的有效性:即使少量注入弱标注的真实缺陷,也能显著提升模型泛化性能;从纯合成缺陷开始训练,可高效融入后续新出现的真实缺陷样本。结果表明,无模型合成缺陷可作为冷启动基线,少量真实缺陷则能优化决策边界以适应未见缺陷特征。提供的无监督与有监督数据划分,旨在推动无监督、自监督及有监督方法在工业场景中的应用。

原文摘要 · Abstract (English)

Automatic visual inspection using machine learning plays a key role in achieving zero-defect policies in industry. Research on anomaly detection is constrained by the availability of datasets that capture complex defect appearances and imperfect imaging conditions, which are typical of production processes. Recent benchmarks indicate that most publicly available datasets are biased towards optimal imaging conditions, leading to an overestimation of their applicability in real-world industrial scenarios. To address this gap, we introduce the Industrial Screen Printing Anomaly Detection Dataset (ISP-AD). It presents challenging small and weakly contrasted surface defects embedded within structured patterns exhibiting high permitted design variability. To the best of our knowledge, it is the largest publicly available industrial dataset to date, including both synthetic and real defects collected directly from the factory floor. Beyond benchmarking recent unsupervised anomaly detection methods, experiments on a mixed supervised training strategy, incorporating both synthesized and real defects, were conducted. Experiments show that even a small amount of injected, weakly labeled real defects improves generalization. Furthermore, starting from training on purely synthetic defects, emerging real defective samples can be efficiently integrated into subsequent scalable training. Overall, our findings indicate that model-free synthetic defects can provide a cold-start baseline, whereas a small number of injected real defects refine the decision boundary for previously unseen defect characteristics. The presented unsupervised and supervised dataset splits are designed to emphasize research on unsupervised, self-supervised, and supervised approaches, enhancing their applicability to industrial settings.

异常检测工业视觉数据集合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。