arXiv:2505.22039cs.CV2025-05被引 19

用多模态推理实现工业缺陷的检测与详细分析

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

  • 通过文本驱动的视觉检测与图文联合推理,无需人工设定阈值
  • 在MMAD基准上达79.1分,超越Qwen2.5-VL-7B和GPT-4o
  • 适合工业质检场景,尤其适用于少样本情况

尽管异常检测已取得显著进展,但生成融合工业知识的详细分析仍具挑战。为此,我们提出OmniAD,一个统一异常检测与理解的新型框架。OmniAD是一种多模态推理模型,结合视觉与文本推理。视觉推理利用Text-as-Mask Encoding进行文本生成驱动的异常检测,无需手动设定阈值;随后,视觉引导的文本推理整合视觉感知,完成全面分析。为提升少样本泛化能力,采用监督微调(SFT)与强化学习(GRPO)相结合的联合训练策略,引入三个精细设计的奖励函数。实验表明,OmniAD在MMAD基准上取得79.1分,优于Qwen2.5-VL-7B和GPT-4o,且在多个异常检测基准上表现优异。结果凸显了增强视觉感知对有效推理的重要性。所有代码与模型将公开。

原文摘要 · Abstract (English)

While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce OmniAD, a novel framework that unifies anomaly detection and understanding for fine-grained analysis. OmniAD is a multimodal reasoner that combines visual and textual reasoning processes. The visual reasoning provides detailed inspection by leveraging Text-as-Mask Encoding to perform anomaly detection through text generation without manually selected thresholds. Following this, Visual Guided Textual Reasoning conducts comprehensive analysis by integrating visual perception. To enhance few-shot generalization, we employ an integrated training strategy that combines supervised fine-tuning (SFT) with reinforcement learning (GRPO), incorporating three sophisticated reward functions. Experimental results demonstrate that OmniAD achieves a performance of 79.1 on the MMAD benchmark, surpassing models such as Qwen2.5-VL-7B and GPT-4o. It also shows strong results across multiple anomaly detection benchmarks. These results highlight the importance of enhancing visual perception for effective reasoning in anomaly understanding. All codes and models will be publicly available.

异常检测多模态工业质检视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。