用视觉上下文对比让大模型精准识别工业缺陷,性能超人类专家。
AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison
- 通过图像对的跨注意力机制增强细粒度视觉对比能力。
- 在MMAD基准上达82.3%准确率,缺陷定位提升3.35倍。
- 专为工业检测设计,适合制造业质检场景应用。
多模态大模型在自然图像理解中表现优异,但在工业异常检测(IAD)中持续表现不佳,主要因训练数据与工业图像差异大,且模型独立编码每张图,仅能在语言空间比较,对细微视觉差异不敏感。为此,我们提出AD-Copilot,一种通过视觉上下文对比实现工业异常检测的交互式多模态大模型。首先构建新颖的数据清洗流程,从稀疏标注的工业图像中挖掘检测知识,生成用于描述、问答和缺陷定位的大规模多模态数据集Chat-AD,富含语义信号。在此基础上,AD-Copilot引入新型对比编码器,利用图像对特征间的交叉注意力增强多图细粒度感知,并采用分阶段策略融合领域知识,逐步提升异常检测能力。此外,我们提出扩展基准MMAD-BBox,以边界框形式评估异常定位。实验表明,AD-Copilot在MMAD基准上达到82.3%准确率,优于所有无数据泄露模型;在MMAD-BBox测试中,最大提升达3.35倍。其性能提升在多个专用与通用基准上具有良好泛化性。尤为显著的是,它在若干IAD任务上超越人类专家水平,展现出作为真实工业检测助手的潜力。所有数据集与模型将开源共享。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform in industrial anomaly detection (IAD). This is because MLLMs trained mostly on general web data differ significantly from industrial images. Moreover, they encode each image independently and can only compare images in the language space, making them insensitive to subtle visual differences that are key to IAD. To tackle these issues, we present AD-Copilot, an interactive MLLM specialized for IAD via visual in-context comparison. We first design a novel data curation pipeline to mine inspection knowledge from sparsely labeled industrial images and generate precise samples for captioning, VQA, and defect localization, yielding a large-scale multimodal dataset Chat-AD rich in semantic signals for IAD. On this foundation, AD-Copilot incorporates a novel Comparison Encoder that employs cross-attention between paired image features to enhance multi-image fine-grained perception, and is trained with a multi-stage strategy that incorporates domain knowledge and gradually enhances IAD skills. In addition, we introduce MMAD-BBox, an extended benchmark for anomaly localization with bounding-box-based evaluation. The experiments show that AD-Copilot achieves 82.3% accuracy on the MMAD benchmark, outperforming all other models without any data leakage. In the MMAD-BBox test, it achieves a maximum improvement of $3.35\times$ over the baseline. AD-Copilot also exhibits excellent generalization of its performance gains across other specialized and general-purpose benchmarks. Remarkably, AD-Copilot surpasses human expert-level performance on several IAD tasks, demonstrating its potential as a reliable assistant for real-world industrial inspection. All datasets and models will be released for the broader benefit of the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。