工业表面缺陷检测框架,兼顾可解释性与鲁棒性评估。
RobustDefect-LLM: Explainable and Robustness-Aware Industrial Surface Defect Classification with Decision Support and AI-Assisted Reporting
- 融合深度学习与可视化证据,支持人工复核的决策流程
- 模型最高准确率99.26%,但严重图像退化下降至40%以下
- 低置信度或置信差小的预测自动转人工,保障关键错误不漏判
本文提出RobustDefect-LLM工业表面缺陷检测框架,整合深度学习分类、面向操作员的视觉证据、置信度感知决策支持、可控的AI辅助报告生成、可追溯存储及移动端交互,形成统一质检工作流。在六类NEU-DET数据集(1,799张图像)上,采用固定划分训练/验证/测试集,评估了四种基于迁移学习的CNN模型:ResNet50、EfficientNet-B0、DenseNet121和MobileNetV3-Large。MobileNetV3-Large表现最佳,测试准确率达99.26%,宏F1为0.9926,95%置信区间为0.9815–1.0000;与DenseNet121相比无显著差异(配对McNemar检验,p=1.000)。单次前向推理平均耗时0.060秒(16.66 FPS)。在合成退化组合下,轻度退化时准确率降至87.78%,强退化时低于40%,显示对图像质量下降敏感。利用Grad-CAM提供视觉证据,置信度低于0.90或前两名预测差距小于0.10的样本转入人工审查。该保守策略实现12.22%的自动覆盖,33个有效案例中观察到选择性准确率为100%(95% CI: 89.43%-100.00%),且所有错误分类均被路由至人工。在正常条件下,100份生成报告全部通过确定性一致性检查,平均延迟1.66秒。结果表明该集成流程可行,但需持续校准与真实场景验证。
原文摘要 · Abstract (English)
This paper presents RobustDefect-LLM, an industrial surface-defect inspection framework integrating deep-learning classification, operator-facing visual evidence, confidence-aware decision support, controlled AI-assisted reporting, traceable storage, and mobile interaction in a unified quality-control workflow. Here, robustness-aware denotes explicit evaluation under controlled image degradation and confidence-aware review routing, not an intrinsic robustness guarantee. Four transfer-learning-based convolutional neural networks, ResNet50, EfficientNet-B0, DenseNet121, and MobileNetV3-Large, were evaluated on 1,799 images from the six-class NEU-DET dataset using fixed training, validation, and held-out in-domain test partitions. MobileNetV3-Large achieved the highest numerical test accuracy (99.26%) and macro F1-score (0.9926), with a bootstrap 95% accuracy CI of 0.9815-1.0000. An exact paired McNemar test found no significant difference from DenseNet121 (p = 1.000). The selected model averaged 0.060 s per CPU forward pass (16.66 FPS). Under combined synthetic degradation, accuracy fell to 87.78% at mild intensity and below 40% at stronger intensities, revealing sensitivity to severe image-quality deterioration. Grad-CAM supplied visual evidence, while predictions with confidence below 0.90 or a top-2 margin below 0.10 were routed to HUMAN REVIEW. This conservative policy provided 12.22% automatic coverage and 100% observed selective accuracy among 33 eligible cases (95% CI: 89.43%-100.00%), while routing both observed classification errors to review. Under nominal controlled conditions, all 100 generated reports passed deterministic consistency checks, with a mean latency of 1.66 s. Results support the feasibility of the integrated workflow while emphasizing the need for calibration, repeated evaluation, and real-world industrial validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。