让AI像工人一样反复检查,自动找证据识别工业缺陷。
AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation
- 用智能体迭代检查,动态调用视觉与外部知识记忆
- 在MMAD数据集上准确率提升5.92%,优于当前最好方法
- 适合需要可解释性检测的工业质检场景
工业异常检测因缺陷细微且局部性强,单次遍历的视觉语言模型常难以捕捉。现有方法缺乏推理时主动获取补充证据的机制。我们提出AgentIAD,一种通过统一动作空间实现迭代工业检测的智能体框架。该智能体在检测过程中动态访问两种记忆:通过感知缩放器(PZ)获取视觉记忆以进行细粒度局部分析,通过网络搜索器(WS)和对比检索器(CR)获取外部知识与跨实例验证。这种设计使模型能通过多轮感知-动作推理逐步积累证据。为在稀疏监督下有效学习此类策略,AgentIAD采用两阶段训练:先进行工具感知的监督微调以初始化结构化推理与记忆访问行为,再通过智能体强化学习优化长程决策策略。大量实验表明,在相同主干网络下,AgentIAD在MMAD基准上分类准确率较之前最优方法提升5.92%,并提供更可靠、可解释的异常分析。
原文摘要 · Abstract (English)
Industrial anomaly detection (IAD) is challenging due to the subtle and highly localized nature of many defects, which single-pass vision--language models (VLMs) often fail to capture. Moreover, existing approaches lack mechanisms to actively acquire complementary evidence during inference. We propose AgentIAD, an agentic vision--language framework that enables iterative industrial inspection through a unified action space. The agent dynamically accesses two forms of memory during inspection: visual memory via the Perceptive Zoomer (PZ) for fine-grained local analysis, and retrieved memory via the Web Searcher (WS) and Comparative Retriever (CR) for external knowledge acquisition and cross-instance verification. This design allows the model to progressively gather evidence through multi-round perception--action reasoning. To effectively learn such policies under sparse supervision, AgentIAD adopts a two-stage training strategy: tool-aware supervised fine-tuning first initializes structured reasoning and memory-access behaviors, followed by agentic reinforcement learning to refine long-horizon decision policies. Extensive experiments show that, under the same backbone, AgentIAD improves classification accuracy by 5.92% over the previous state-of-the-art method on the MMAD benchmark while providing more reliable and interpretable anomaly analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。