用堆叠提示增强图文对齐,提升工业缺陷检测的零样本性能
StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection
- 通过聚类驱动的堆叠提示生成通用语义提示,提升特征区分度
- 在7个数据集上实现领先性能,异常分割准确率显著提升
- 适合需要快速适配新类别、追求高泛化能力的工业质检场景
在零样本工业缺陷检测中,提升CLIP模型中文本与图像特征的对齐程度是关键挑战。现有方法多依赖特定类别提示预训练,易导致过拟合并限制泛化能力。为此,我们提出StackCLIP模型,通过多类别名称堆叠生成堆叠提示。核心包含两个组件:聚类驱动堆叠提示(CSP)模块利用语义相近类别堆叠生成通用提示,并通过多对象文本特征融合增强相似物体间的异常区分性;集成特征对齐(EFA)模块为每个堆叠聚类训练专用线性层,并根据测试类别属性自适应融合。该框架显著提升训练速度、稳定性和收敛性,大幅改善异常分割表现。此外,引入调节提示学习(RPL)模块,利用堆叠提示的泛化能力优化提示学习,进一步提升分类任务性能。在7个工业缺陷检测数据集上的实验证明,该方法在零样本异常检测与分割任务中均达到当前最优水平。
原文摘要 · Abstract (English)
Enhancing the alignment between text and image features in the CLIP model is a critical challenge in zero-shot industrial anomaly detection tasks. Recent studies predominantly utilize specific category prompts during pretraining, which can cause overfitting to the training categories and limit model generalization. To address this, we propose a method that transforms category names through multicategory name stacking to create stacked prompts, forming the basis of our StackCLIP model. Our approach introduces two key components. The Clustering-Driven Stacked Prompts (CSP) module constructs generic prompts by stacking semantically analogous categories, while utilizing multi-object textual feature fusion to amplify discriminative anomalies among similar objects. The Ensemble Feature Alignment (EFA) module trains knowledge-specific linear layers tailored for each stack cluster and adaptively integrates them based on the attributes of test categories. These modules work together to deliver superior training speed, stability, and convergence, significantly boosting anomaly segmentation performance. Additionally, our stacked prompt framework offers robust generalization across classification tasks. To further improve performance, we introduce the Regulating Prompt Learning (RPL) module, which leverages the generalization power of stacked prompts to refine prompt learning, elevating results in anomaly detection classification tasks. Extensive testing on seven industrial anomaly detection datasets demonstrates that our method achieves state-of-the-art performance in both zero-shot anomaly detection and segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。