arXiv:2507.21619cs.CV2025-07被引 11

用难例感知策略提升大模型工业缺陷检测能力

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

  • 基于难例感知的强化学习优化,动态调整模型学习重点
  • 在7个任务上使基线模型平均性能提升7.77%
  • 适合需要少样本缺陷检测的工业场景应用

工业异常检测对保障制造系统安全可靠至关重要。尽管多模态大语言模型具备强大的视觉-语言推理能力,但缺乏领域适配时其在工业异常检测中的效果仍有限。本文提出EMIT框架,通过难例感知的组相对策略优化(GRPO)增强多模态大模型在工业异常检测中的表现。构建了多任务工业异常检测数据集,并利用GPT生成的对象文本描述弥补缺陷图像缺失问题。针对少样本异常检测,融合软提示与基于补丁级对比的热图引导嵌入。为更好处理模型难以判断的困难样本,提出难例感知的GRPO:引入响应重采样策略确保正确答案被采样,结合优势重加权机制强化对困难样本的学习。在MMAD基准上的大量实验表明,EMIT显著提升多模态大模型在工业异常检测中的性能,相较基线模型InternVL3-8B,在7个任务上平均提升7.77%。

原文摘要 · Abstract (English)

Industrial anomaly detection (IAD) plays a crucial role in maintaining the safety and reliability of manufacturing systems. While multimodal large language models (MLLMs) show strong vision-language reasoning abilities, their effectiveness in IAD remains limited without domain-specific adaptation. In this work, we propose EMIT, a unified framework that enhances MLLMs for IAD via difficulty-aware group relative policy optimization (GRPO). EMIT constructs a multi-task IAD dataset and utilizes GPT-generated object text descriptions to compensate for missing defective images. For few-shot anomaly detection, it integrates a soft prompt and heatmap-guided contrastive embeddings derived from patch-level comparisons. To better handle difficult data samples, i.e., cases where the MLLM struggles to generate correct answers, we propose a difficulty-aware GRPO that extends the original GRPO by incorporating a response resampling strategy to ensure the inclusion of correct answers in the sampled responses, as well as an advantage reweighting mechanism to strengthen learning from such difficult data samples. Extensive experiments on the MMAD benchmark demonstrate that EMIT significantly enhances the IAD performance of MLLMs, achieving an average improvement of 7.77\% over the base model (InternVL3-8B) across seven tasks.

工业检测多模态模型强化学习少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。