arXiv:2604.17949cs.CV2026-04

让工业缺陷检测可解释,用多模态生成带证据的异常报告

ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection

论文配图:ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection
图 1 · 摘自论文原文
  • 用语言引导双跳定位,从图像传感器中提取异常线索
  • 零样本下实现像素级异常掩码,准确率超越已有方法
  • 适合需要可信检测结果的制造业质量控制场景

基于深度学习的工业异常检测器通常为黑箱,难以提供具有物理意义的缺陷证据。本文提出ZSG-IAD,一种多模态视觉-语言框架,用于零样本地面化工业异常检测。输入包括RGB图像、传感器图像和3D点云,ZSG-IAD生成结构化异常报告与像素级异常掩码。该框架引入语言引导的双跳定位模块:(1) 异常相关语句从多模态特征中选择类似证据的潜在槽位,获得粗粒度空间支持;(2) 选定槽位通过通道-空间门控与轻量解码器调制特征图,生成细粒度掩码。为提升可靠性,进一步采用可执行规则强化学习(Executable-Rule GRPO),以可验证奖励促进结构化输出、异常区域一致性和推理-结论连贯性。在多个工业异常检测基准上的实验表明,ZSG-IAD在零样本设置下表现优异,且解释更透明、更具物理依据。代码与标注将公开,以支持可信工业异常检测系统的后续研究。

原文摘要 · Abstract (English)

Deep learning-based industrial anomaly detectors often behave as black boxes, making it hard to justify decisions with physically meaningful defect evidence. We propose ZSG-IAD, a multimodal vision-language framework for zero-shot grounded industrial anomaly detection. Given RGB images, sensor images, and 3D point clouds, ZSG-IAD generates structured anomaly reports and pixel-level anomaly masks. ZSG-IAD introduces a language-guided two-hop grounding module: (1) anomaly-related sentences select evidence-like latent slots distilled from multimodal features, yielding coarse spatial support; (2) selected slots modulate feature maps via channel-spatial gating and a lightweight decoder to produce fine-grained masks. To improve reliability, we further apply Executable-Rule GRPO with verifiable rewards to promote structured outputs, anomaly-region consistency, and reasoning-conclusion coherence. Experiments across multiple industrial anomaly benchmarks show strong zero-shot performance and more transparent, physically grounded explanations than prior methods. We will release code and annotations to support future research on trustworthy industrial anomaly detection systems.

异常检测多模态可解释性工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。