构建了工业缺陷检测最大最多样数据集,推动模型泛化能力研究
Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era

- 构建160类、近20万张高分辨率图像的跨行业多材料缺陷数据集
- 主流无监督模型在类别数从30增至160时性能下降10%-20%
- 零样本与小样本模型表现稳定,适合工业场景泛化研究
工业异常检测(IAD)是保障生产安全、提升产品质量和优化制造效率的核心。然而现有公开基准数据集存在类别多样性不足、规模有限的问题,导致算法性能饱和且难以迁移至复杂真实场景。为此,我们提出Real-IAD Variety,目前规模最大、最多样化的IAD基准数据集,包含198,950张高分辨率图像,覆盖160个物体类别。其多样性体现在28个行业、24种材料类型、22种颜色变化及27种缺陷类型。实验表明,当前最先进的多类无监督异常检测方法在类别数从30增至160时性能下降10%至20%;而零样本与少样本模型展现出强鲁棒性,性能稳定,显著提升跨工业场景泛化能力。该数据集为下一代基础型IAD模型的训练与评估提供了关键资源。
原文摘要 · Abstract (English)
Industrial Anomaly Detection (IAD) is a cornerstone for ensuring operational safety, maintaining product quality, and optimizing manufacturing efficiency. However, the advancement of IAD algorithms is severely hindered by the limitations of existing public benchmarks. Current datasets often suffer from restricted category diversity and insufficient scale, leading to performance saturation and poor model transferability in complex, real-world scenarios. To bridge this gap, we introduce Real-IAD Variety, the largest and most diverse IAD benchmark. It comprises 198,950 high-resolution images across 160 distinct object categories. The dataset ensures unprecedented diversity by covering 28 industries, 24 material types, 22 color variations, and 27 defect types. Our extensive experimental analysis highlights the substantial challenges posed by this benchmark: state-of-the-art multi-class unsupervised anomaly detection methods suffer significant performance degradation (ranging from 10% to 20%) when scaled from 30 to 160 categories. Conversely, we demonstrate that zero-shot and few-shot IAD models exhibit remarkable robustness to category scale-up, maintaining consistent performance and significantly enhancing generalization across diverse industrial contexts. This unprecedented scale positions Real-IAD Variety as an essential resource for training and evaluating next-generation foundation IAD models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。