实测19种异常检测模型在工业场景中表现不稳定,提出人机协同框架提升缺陷检测可靠性。
From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection

- 构建人机协同框架,融合标注、AI辅助与验证引擎
- 在BowTie数据集上发现模型性能受预处理影响大,无统一最优方案
- 适合工业质检人员、部署工程师及关注落地的AI研究者
自动化异常检测方法在学术基准上表现优异,但在真实工业环境中表现不明。本文在包含反光表面、细微缺陷和轮廓特异性变化的BowTie制造数据集上评估了19种无监督异常检测模型。结果表明,相比标准基准(如MVTec AD),模型性能更不稳定,对预处理高度敏感,且在不同条件下不一致,未出现普遍鲁棒的方法;共识审计进一步显示数据质量直接影响部署效果。基于此,我们开发并初步部署了一个统一的人机协同框架,用于工件检测,替代原有的手动视觉检查与文档流程。该系统支持热图引导的缺陷审查、SAM优化的候选区域调整、掩码评估以及审查历史记录,以确保检查一致性与新人培训。结果揭示了基准性能与实际部署之间的差距,并提供了可行的应对方案。
原文摘要 · Abstract (English)
Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial conditions is less clear. In this work, we evaluate 19 unsupervised anomaly detection models on the BowTie dataset, a challenging manufacturing dataset with reflective surfaces, subtle defects, and profile-specific variation. In contrast to benchmark results, we observe that model performance is less stable than typically reported on standard benchmarks such as MVTec AD, highly sensitive to preprocessing, and inconsistent across conditions, with no single approach emerging as uniformly robust; a consensus audit further indicates that nominal-data quality affects deployment. Motivated by these findings, we developed and initially deployed a unified human-in-the-loop framework for manufactured-part inspection that combines image annotation, AI-assisted defect detection, and an integrated validation engine, replacing a prior manual visual inspection and documentation workflow. The system supports heatmap-guided defect review, SAM-refined candidate regions for inspector acceptance, rejection, or boundary adjustment, mask evaluation where annotations exist, and review history for inspector consistency and onboarding. Together, the results highlight the gap between benchmark performance and deployment reality, and provide a practical framework for addressing it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。