arXiv:2508.04175cs.CV2025-08被引 8

用多阶段推理和细粒度奖励优化,让大模型更懂工业缺陷检测。

AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization

  • 分步骤引导模型从定位到精细分析,生成多样化回答
  • 用连续奖励信号替代二元反馈,提升判断准确性
  • 适合需要精准识别微小制造缺陷的工业场景

尽管多模态大语言模型在多个领域表现出色,但在专门的异常检测任务中仍受限于领域适应难题。现有基于组相对策略优化(GRPO)的方法存在两大缺陷:模型产生一致回复时训练数据利用不足,且缺乏对推理过程的充分监督,导致过早做出二分类决策而缺少深度分析。本文提出一个综合框架,通过两项协同创新解决上述问题。首先,引入多阶段推理性流程,引导模型从区域识别逐步过渡到聚焦检查,生成多样化的响应模式,有利于GRPO优化,并实现对分析流程的结构化监督。其次,设计细粒度奖励机制,融合分类准确率与定位监督,将二元反馈转化为连续信号,区分真实分析洞察与偶然正确。在多个工业数据集上的全面评估表明,该方法显著提升了通用视觉-语言模型向专业异常检测任务的适配性能。所提方法在保持高效标注利用的前提下,实现了更高精度,有效弥合了通用多模态大模型能力与检测细微制造缺陷及结构异常所需的精细化视觉判别之间的差距。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrained by domain adaptation challenges. Existing Group Relative Policy Optimization (GRPO) based approaches suffer from two critical limitations: inadequate training data utilization when models produce uniform responses, and insufficient supervision over reasoning processes that encourage immediate binary decisions without deliberative analysis. We propose a comprehensive framework addressing these limitations through two synergistic innovations. First, we introduce a multi-stage deliberative reasoning process that guides models from region identification to focused examination, generating diverse response patterns essential for GRPO optimization while enabling structured supervision over analytical workflows. Second, we develop a fine-grained reward mechanism incorporating classification accuracy and localization supervision, transforming binary feedback into continuous signals that distinguish genuine analytical insight from spurious correctness. Comprehensive evaluation across multiple industrial datasets demonstrates substantial performance improvements in adapting general vision-language models to specialized anomaly detection. Our method achieves superior accuracy with efficient adaptation of existing annotations, effectively bridging the gap between general-purpose MLLM capabilities and the fine-grained visual discrimination required for detecting subtle manufacturing defects and structural irregularities.

异常检测多模态大模型推理优化工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。