arXiv:2509.25856cs.CV2025-09中稿 · ed被引 1

提出统一的纯图像提示框架,实现无需训练的工业缺陷检测

PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection

  • 构建纯视觉提示机制,包含对齐模块和前景掩码
  • 少样本与零样本场景下性能超越现有方法
  • 适合快速部署且无需文本提示的工业质检场景

工业异常检测正越来越多依赖基础模型,以实现强分布外泛化能力与实际部署中的快速适应。以往研究多集中于文本提示调优,而视觉提示部分则因各基础模型处理流程差异导致碎片化。本文提出统一的基于补丁的异常检测框架——PatchEAD,支持无需训练的异常检测,兼容多种基础模型。该框架通过构建视觉提示技术,包括对齐模块与前景掩码,实现了在无文本特征情况下的优异表现。实验表明,其在少样本及批量零样本设置下均优于现有方法。研究还分析了骨干结构与预训练特性对补丁相似性鲁棒性的影响,为真实视觉检测中基础模型的选择与配置提供可操作指导。结果表明,一个良好统一的纯补丁框架可实现无需校准的快速部署,避免复杂文本提示工程。

原文摘要 · Abstract (English)

Industrial anomaly detection is increasingly relying on foundation models, aiming for strong out-of-distribution generalization and rapid adaptation in real-world deployments. Notably, past studies have primarily focused on textual prompt tuning, leaving the intrinsic visual counterpart fragmented into processing steps specific to each foundation model. We aim to address this limitation by proposing a unified patch-focused framework, Patch-Exclusive Anomaly Detection (PatchEAD), enabling training-free anomaly detection that is compatible with diverse foundation models. The framework constructs visual prompting techniques, including an alignment module and foreground masking. Our experiments show superior few-shot and batch zero-shot performance compared to prior work, despite the absence of textual features. Our study further examines how backbone structure and pretrained characteristics affect patch-similarity robustness, providing actionable guidance for selecting and configuring foundation models for real-world visual inspection. These results confirm that a well-unified patch-only framework can enable quick, calibration-light deployment without the need for carefully engineered textual prompts.

异常检测视觉提示工业质检零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。