arXiv:2509.20479cs.CV2025-09被引 1

实测发现大模型在真实工业缺陷检测中表现不佳,远不如公开数据集上理想。

Are Foundation Models Ready for Industrial Defect Recognition? A Reality Check on Real-World Data

  • 用文本提示直接调用大模型进行缺陷识别,免去繁琐标注
  • 所有测试大模型在真实工业图像上全失败,准确率远低于预期
  • 适合关注工业落地挑战的研究者和工程师参考

基础模型(FMs)在文本和图像处理任务中表现出色,具备零样本跨域泛化能力。这使其在量产过程中的自动质量检测中具有潜力——仅通过文本描述异常即可评估多种产品图像,避免为每个产品单独标注训练数据。然而,我们在自建的真实工业图像数据及公开数据集上测试多个近期基础模型,发现它们在真实数据上全部失效,而在公共基准数据集上却表现良好。结果表明,当前基础模型尚未准备好应对真实世界复杂多变的工业场景。

原文摘要 · Abstract (English)

Foundation Models (FMs) have shown impressive performance on various text and image processing tasks. They can generalize across domains and datasets in a zero-shot setting. This could make them suitable for automated quality inspection during series manufacturing, where various types of images are being evaluated for many different products. Replacing tedious labeling tasks with a simple text prompt to describe anomalies and utilizing the same models across many products would save significant efforts during model setup and implementation. This is a strong advantage over supervised Artificial Intelligence (AI) models, which are trained for individual applications and require labeled training data. We test multiple recent FMs on both custom real-world industrial image data and public image data. We show that all of those models fail on our real-world data, while the very same models perform well on public benchmark datasets.

缺陷检测工业视觉大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。