用多专家框架让大模型更懂工业缺陷检测。
Can Multimodal Large Language Models be Guided to Improve Industrial Anomaly Detection?
- 设计四模块专家系统,融合上下文参考与领域知识。
- 在MMAD数据集上精度与鲁棒性显著提升。
- 适合需要高精度缺陷识别的工业场景使用。
工业环境中,准确检测异常对保障产品质量和运行安全至关重要。传统工业异常检测(IAD)模型在动态生产环境下灵活性与适应性不足,难以应对新缺陷类型和工艺变化。多模态大语言模型(MLLM)虽具备强大的通用视觉理解能力,但缺乏行业特定知识(如缺陷容忍度),限制了其在IAD中的应用。为此,本文提出Echo——一种多专家框架,包含参考提取器(检索相似正常图像作为基准)、知识引导模块(提供领域知识)、推理专家(支持结构化分步推理)和决策模块(整合各模块输出,生成精准上下文响应)。在MMAD基准测试中,Echo显著提升了模型的适应性、精度与鲁棒性,更贴近真实工业场景需求。
原文摘要 · Abstract (English)
In industrial settings, the accurate detection of anomalies is essential for maintaining product quality and ensuring operational safety. Traditional industrial anomaly detection (IAD) models often struggle with flexibility and adaptability, especially in dynamic production environments where new defect types and operational changes frequently arise. Recent advancements in Multimodal Large Language Models (MLLMs) hold promise for overcoming these limitations by combining visual and textual information processing capabilities. MLLMs excel in general visual understanding due to their training on large, diverse datasets, but they lack domain-specific knowledge, such as industry-specific defect tolerance levels, which limits their effectiveness in IAD tasks. To address these challenges, we propose Echo, a novel multi-expert framework designed to enhance MLLM performance for IAD. Echo integrates four expert modules: Reference Extractor which provides a contextual baseline by retrieving similar normal images, Knowledge Guide which supplies domain-specific insights, Reasoning Expert which enables structured, stepwise reasoning for complex queries, and Decision Maker which synthesizes information from all modules to deliver precise, context-aware responses. Evaluated on the MMAD benchmark, Echo demonstrates significant improvements in adaptability, precision, and robustness, moving closer to meeting the demands of real-world industrial anomaly detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。