JUDO通过对比正常与缺陷图像,提升工业异常检测的精准推理能力。
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA

- 用正常图对比缺陷图,实现细粒度视觉差异定位
- 结合监督微调与强化学习,注入领域知识提升推理准确率
- 适合需要高精度工业质检的场景,如制造业质量控制
工业异常检测得益于大视觉语言模型(LMMs)的发展,可支持多样化的人类指令,尤其在基于视觉的推理方面提升了图像理解能力。然而,现有LMMs缺乏领域特定知识,难以在复杂工业场景中生成准确回答。本文提出JUDO——一种并置式领域导向多模态推理框架,通过将查询图像与正常图像对比作为视觉上下文,实现细粒度缺陷区域分割;同时通过监督微调注入领域知识,并采用定制奖励的强化学习(GRPO)引导领域推理过程。实验表明,JUDO在MMAD基准上超越Qwen2.5-VL-7B和GPT-4o等模型,验证了增强领域知识与上下文对异常理解推理的关键作用。
原文摘要 · Abstract (English)
Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, particularly through visually grounded reasoning for better image understanding. However, LMMs lack domain-specific knowledge, which limits their ability to generate accurate responses in complex industrial scenarios. In this work, we present JUDO, Juxtaposed Domain-Oriented Multimodal Reasoner, a framework that efficiently incorporates domain knowledge and context in visual and textual reasoning. Through visual reasoning, our model segments the defect region by juxtaposing query images with normal images as visual domain context, enabling a fine-grained visual comparative inspection. Furthermore, we inject domain knowledge through supervised fine-tuning (SFT) to enhance context understanding and subsequently guide domain reasoning through reinforcement learning (GRPO) with tailored rewards, opting for a domain-oriented reasoning process. Experimental results demonstrate that JUDO achieves superior performance on the MMAD benchmark, surpassing models such as Qwen2.5-VL-7B and GPT-4o. These results highlight the importance of enhancing domain knowledge and context for effective reasoning in anomaly understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。