发现视觉语言模型中隐藏的异常敏感神经元,无需训练即可提升检测性能。
Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

- 通过少量正常样本挖掘预训练模型中的潜在异常知识
- 在工业异常检测基准上达到顶尖性能,且仅需极小计算开销
- 适用于追求高效、可解释异常检测的工业场景
大规模视觉语言模型(VLMs)展现出强大的零样本能力,但其异常检测(AD)性能的内在机制仍不明确。现有方法多将VLM视为黑箱特征提取器,认为异常知识需通过外部适配器或记忆库获取。本文挑战这一假设,提出异常知识本质上内嵌于预训练模型中,但处于隐匿状态且激活不足。我们假设这些知识集中于少数异常敏感神经元。为此,提出无需训练的潜藏异常知识挖掘框架LAKE,仅用少量正常样本即可识别并激发这些关键神经信号。通过分离这些敏感神经元,LAKE构建了高度紧凑的正常性表征,融合视觉结构偏差与跨模态语义激活。在工业异常检测基准上的大量实验表明,LAKE实现当前最优性能,并提供神经元级别的可解释性。本工作倡导范式转变:将异常检测重新定义为对预训练模型中潜藏知识的定向激活,而非下游任务的知识获取。
原文摘要 · Abstract (English)
Large-scale vision-language models (VLMs) exhibit remarkable zero-shot capabilities, yet the internal mechanisms driving their anomaly detection (AD) performance remain poorly understood. Current methods predominantly treat VLMs as black-box feature extractors, assuming that anomaly-specific knowledge must be acquired through external adapters or memory banks. In this paper, we challenge this assumption by arguing that anomaly knowledge is intrinsically embedded within pre-trained models but remains latent and under-activated. We hypothesize that this knowledge is concentrated within a sparse subset of anomaly-sensitive neurons. To validate this, we propose latent anomaly knowledge excavation (LAKE), a training-free framework that identifies and elicits these critical neuronal signals using only a minimal set of normal samples. By isolating these sensitive neurons, LAKE constructs a highly compact normality representation that integrates visual structural deviations with cross-modal semantic activations. Extensive experiments on industrial AD benchmarks demonstrate that LAKE achieves state-of-the-art performance while providing intrinsic, neuron-level interpretability. Ultimately, our work advocates for a paradigm shift: redefining anomaly detection as the targeted activation of latent pre-trained knowledge rather than the acquisition of a downstream task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。