用概念引导动态注意力,提升少样本工业缺陷检测精度
ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection

- 基于基础模型特征,通过固定正常概念重构查询特征统计
- 在多数据集上优于现有方法,不同样本量下均显著提效
- 轻量设计适合快速推理,适合工业质检场景落地
少样本工业异常检测(FS-IAD)旨在冷启动阶段仅凭少量正常样本识别和定位工业图像中的视觉缺陷。尽管当前方法依赖基础模型特征并取得良好表现,但因正常样本极度稀缺,模型泛化能力仍较脆弱。为此,我们提出ConceptADapt:一种概念引导的自适应特征重建模型,结合动态注意力机制。该模型从有限支持特征中预学习一组固定正常概念,并利用其与查询特征的关联关系,重新校准特征统计以提升测试时检测性能。为缓解低数据环境下常见的特征捷径问题,我们引入融合稀疏自编码器的动态注意力机制,训练中学习鲁棒正常概念。此外,为实现推理快速适配,模型在注意力模块中采用LoRA,仅引入极少量可更新参数。在MVTec-AD、VisA和MPDD三个主流基准上的大量实验表明,本模型在检测与定位任务中持续优于当前最优方法,且在多种样本设置下均有显著提升。
原文摘要 · Abstract (English)
Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase, where only a limited number of normal training samples are available per category. Recent advances in this field predominantly leverage visual features from foundation-model and have achieved promising performance. Despite the strong representational power of foundation-model features, the model generalization remains fragile due to the extreme scarcity of normal training data.To address this pivotal issue, we propose ConceptADapt, a concept-guided adaptive feature reconstruction model with dynamic attention. Specifically, our model pre-learns a set of fixed normal concepts from the limited support features and leverages them to mine relationships with query features, thereby recalibrating their statistics for improved anomaly detection at test time. To mitigate the prevalent feature shortcut problem, which is particularly severe under low-data regimes, we further develop a dynamic attention mechanism integrated with sparse autoencoders to learn robust normal concepts during training. Moreover, to enable fast adaptation during inference, our model remains lightweight by incorporating LoRA into the attention module, which introduces only minimal updating parameters.Extensive experiments on three widely adopted FS-IAD benchmarks, including MVTec-AD, VisA, and MPDD, demonstrate that our model consistently outperforms state-of-the-art (SOTA) approaches across both detection and localization tasks, achieving significant improvements under various shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。