arXiv:2607.18850cs.CVcs.AI2026-07被引 1

用语言判断指导像素级缺陷定位,提升工业异常检测精度。

OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation

论文配图:OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation
图 1 · 摘自论文原文
  • 通过自蒸馏框架将语言判断转化为视觉引导信号
  • 在多个指标上超越现有基于视觉语言模型的方法
  • 适合需要精准缺陷定位的工业质检场景

大型视觉语言模型(LVLM)在工业异常检测(IAD)中展现出强大潜力,能提供图像级异常判断与可解释的缺陷推理。然而,现有方法难以从生成的语言判断中产出精确的像素级异常图。本文提出OPD-IAD——一种基于证据优先的密集在线自蒸馏框架,将特权缺陷证据蒸馏至模型自身的在线判断轨迹中,使最终判断在密集监督下学习,而非仅作为文本答案处理。该判断作为语义条件,驱动密集异常感知。为此引入语言引导视觉锚定机制:通过重推理将图像与问题在最终判断条件下重新编码为语义锚点,并通过对比热图头与密集视觉特征进行对比,生成异常图。语言判断提供紧凑语义引导,而密集视觉特征仍为像素级评分基础,避免语言质量直接决定像素响应。大量实验表明,OPD-IAD在多数图像级、像素级及问答指标上均达到最优性能。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable defect reasoning. However, current LVLM-based IAD methods still struggle to produce precise pixel-level anomaly maps from generated language judgments. We aim to achieve precise pixel-level localization while using language as guidance rather than letting it dominate the visual response. Specifically, we propose \textbf{OPD-IAD}, an evidence-privileged dense on-policy self-distillation framework for LVLM-based IAD. OPD-IAD distills privileged defect evidence onto the model's own on-policy judgment trajectory, enabling the final generated judgment to be learned under dense supervision rather than treated only as a textual answer. The resulting judgment serves as a semantic condition for dense anomaly perception. To turn this condition into dense visual evidence, we introduce \textbf{Language-guided Visual Anchoring}, which uses a judgment reforward to re-encode the image and question under the final-judgment condition into semantic anchors and contrasts them with dense visual features through a contrastive heatmap head to generate anomaly maps. The language judgment therefore provides compact semantic guidance, while dense visual features remain the basis for pixel-level scoring, allowing language to guide anomaly localization without letting language quality directly dictate the pixel-level response. Extensive experiments show that OPD-IAD achieves the best overall performance among LVLM-based IAD methods, leading on most image-level, pixel-level, and QA metrics.

工业检测视觉语言模型异常定位自蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。