用知识增强推理,让AI看懂工业缺陷并给出精准解释。
ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning
- 构建图文结合的知识库,支持缺陷知识的精准检索与理解。
- 在零样本检测任务中达到当前最优性能,跨类别泛化能力强。
- 适合需要高可靠性缺陷分析的工业视觉检测场景。
工业视觉检测对自动化质检至关重要。尽管多模态大模型具备强大的语言理解能力,其表现仍远逊于人类专家。本文识别出两大挑战:(i) 预训练阶段缺乏异常检测知识整合;(ii) 异常推理时语言生成缺乏技术精确性与上下文感知。为此,我们提出ADSeeker,一个基于知识的异常推理框架。首先,构建了包含语义丰富描述与图像-文档对的视觉文档知识库SEEK-M&V,弥补现有资源仅依赖非结构化文本的不足。为高效检索与利用知识,引入查询图像-知识检索增强生成(Q2K RAG)框架。为提升零样本异常检测(ZSAD)性能,采用分层稀疏提示机制和类型级特征提取异常模式。针对工业异常检测数据稀缺问题,构建了最大规模的异常数据集Multi-type Anomaly MulA,涵盖26个类别中的72种多尺度缺陷。大量实验表明,该即插即用框架在多个基准数据集上实现了领先的零样本检测效果。
原文摘要 · Abstract (English)
Automatic vision inspection holds significant importance in industry inspection. While multimodal large language models (MLLMs) exhibit strong language understanding capabilities and hold promise for this task, their performance remains significantly inferior to that of human experts. In this context, we identify two key challenges: (i) insufficient integration of anomaly detection (AD) knowledge during pre-training, and (ii) the lack of technically precise and context-aware language generation for anomaly reasoning. To address these issues, we propose ADSeeker, an anomaly task assistant designed to enhance inspection performance through knowledge-grounded reasoning. ADSeeker first leverages a curated visual document knowledge base, SEEK-M&V, which we construct to address the limitations of existing resources that rely solely on unstructured text. SEEK-M\&V includes semantic-rich descriptions and image-document pairs, enabling more comprehensive anomaly understanding. To effectively retrieve and utilize this knowledge, we introduce the Query Image-Knowledge Retrieval-Augmented Generation Q2K RAG framework. To further enhance the performance in zero-shot anomaly detection (ZSAD), ADSeeker leverages the Hierarchical Sparse Prompt mechanism and type-level features to efficiently extract anomaly patterns. Furthermore, to tackle the challenge of limited industry anomaly detection (IAD) data, we introduce the largest-scale AD dataset, Multi-type Anomaly MulA, encompassing 72 multi-scale defect types across 26 categories. Extensive experiments show that our plug-and-play framework, ADSeeker, achieves state-of-the-art zero-shot performance on several benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。