用类人方式实现工业缺陷检测与推理的统一框架
IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
- 通过三阶段训练让模型逐步掌握工业知识和异常感知能力
- 仅需少量样本即可泛化到新产品的缺陷检测,性能显著提升
- 支持图像级、像素级评分及语言解释,适合工业质检场景
少样本工业异常检测(FS-IAD)在自动化质量检验中具有重要应用。现有基于大视觉语言模型(LVLM)的方法虽取得一定进展,但普遍缺乏工业领域知识与推理能力,难以媲美人类质检员。为此,我们提出统一框架IADGPT,模拟人类行为实现少样本下的异常检测、定位与推理,适用于多样且新颖的工业产品。采用三阶段渐进式训练:前两阶段逐步构建基础工业知识与差异感知能力;第三阶段引入上下文学习范式,利用少量样本作为示例实现对新产品的良好泛化。此外,设计策略结合模型输出的逻辑值与注意力图,分别生成图像级和像素级异常分数,并配合语言输出完成异常推理。为支持训练,构建包含400类工业产品、10万张图像及丰富属性级文本标注的新数据集。实验表明,IADGPT在异常检测上表现优异,且在定位与推理任务中具竞争力。数据集将于最终版本发布。
原文摘要 · Abstract (English)
Few-Shot Industrial Anomaly Detection (FS-IAD) has important applications in automating industrial quality inspection. Recently, some FS-IAD methods based on Large Vision-Language Models (LVLMs) have been proposed with some achievements through prompt learning or fine-tuning. However, existing LVLMs focus on general tasks but lack basic industrial knowledge and reasoning capabilities related to FS-IAD, making these methods far from specialized human quality inspectors. To address these challenges, we propose a unified framework, IADGPT, designed to perform FS-IAD in a human-like manner, while also handling associated localization and reasoning tasks, even for diverse and novel industrial products. To this end, we introduce a three-stage progressive training strategy inspired by humans. Specifically, the first two stages gradually guide IADGPT in acquiring fundamental industrial knowledge and discrepancy awareness. In the third stage, we design an in-context learning-based training paradigm, enabling IADGPT to leverage a few-shot image as the exemplars for improved generalization to novel products. In addition, we design a strategy that enables IADGPT to output image-level and pixel-level anomaly scores using the logits output and the attention map, respectively, in conjunction with the language output to accomplish anomaly reasoning. To support our training, we present a new dataset comprising 100K images across 400 diverse industrial product categories with extensive attribute-level textual annotations. Experiments indicate IADGPT achieves considerable performance gains in anomaly detection and demonstrates competitiveness in anomaly localization and reasoning. We will release our dataset in camera-ready.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。