构建超大规模物流缺陷检测数据集,挑战现有方法极限
Kaputt: A Large-Scale Dataset for Visual Defect Detection
- 构建包含23万张图像的物流缺陷数据集,覆盖4.8万种不同物品
- 现有先进方法在该数据集上最高仅达56.96% AUROC,显著低于制造场景
- 适合研究物流视觉异常检测、工业质检新范式的研究者使用
我们提出一个面向物流场景的新型大规模缺陷检测数据集。当前工业异常检测主要聚焦于姿态可控、类别有限的制造场景,现有基准如MVTec-AD和VisA已接近性能饱和,顶尖方法AUROC可达99.9%。而零售物流中的异常检测面临更复杂的姿态与外观多样性挑战,现有方法表现明显不足。为此,我们引入新基准:数据集含超过23万张图像(超2.9万例缺陷样本),规模是MVTec-AD的40倍,涵盖超过4.8万种不同物体。通过广泛评估多种先进异常检测方法,发现其在该数据集上最高仅达56.96% AUROC。定性分析表明,现有方法难以在强姿态与外观变化下有效利用正常样本。本数据集为零售物流异常检测设立新基准,推动后续研究发展。数据集可访问 https://www.kaputt-dataset.com。
原文摘要 · Abstract (English)
We present a novel large-scale dataset for defect detection in a logistics setting. Recent work on industrial anomaly detection has primarily focused on manufacturing scenarios with highly controlled poses and a limited number of object categories. Existing benchmarks like MVTec-AD [6] and VisA [33] have reached saturation, with state-of-the-art methods achieving up to 99.9% AUROC scores. In contrast to manufacturing, anomaly detection in retail logistics faces new challenges, particularly in the diversity and variability of object pose and appearance. Leading anomaly detection methods fall short when applied to this new setting. To bridge this gap, we introduce a new benchmark that overcomes the current limitations of existing datasets. With over 230,000 images (and more than 29,000 defective instances), it is 40 times larger than MVTec-AD and contains more than 48,000 distinct objects. To validate the difficulty of the problem, we conduct an extensive evaluation of multiple state-of-the-art anomaly detection methods, demonstrating that they do not surpass 56.96% AUROC on our dataset. Further qualitative analysis confirms that existing methods struggle to leverage normal samples under heavy pose and appearance variation. With our large-scale dataset, we set a new benchmark and encourage future research towards solving this challenging problem in retail logistics anomaly detection. The dataset is available for download under https://www.kaputt-dataset.com.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。