综述大模型与Transformer在视觉异常检测中的应用进展
Foundation Models and Transformers for Anomaly Detection: A Survey
- 按重建、特征、零/少样本分类方法,梳理技术演进
- 大模型结合注意力机制,提升检测鲁棒性与可解释性
- 适合关注AI异常检测前沿的科研与工程人员
随着深度学习的发展,本文综述了Transformer和基础模型在推动视觉异常检测(VAD)方面的变革作用。这些架构凭借全局感受野和强适应性,有效应对长程依赖建模、上下文建模及数据稀缺等挑战。文章将VAD方法分为基于重建、基于特征以及零/少样本三类,突出基础模型带来的范式转变。通过整合注意力机制并利用大规模预训练,Transformer与基础模型实现了更鲁棒、可解释且可扩展的异常检测方案。本工作系统回顾了当前主流技术、其优势与局限,并展望了该领域的发展趋势。
原文摘要 · Abstract (English)
In line with the development of deep learning, this survey examines the transformative role of Transformers and foundation models in advancing visual anomaly detection (VAD). We explore how these architectures, with their global receptive fields and adaptability, address challenges such as long-range dependency modeling, contextual modeling and data scarcity. The survey categorizes VAD methods into reconstruction-based, feature-based and zero/few-shot approaches, highlighting the paradigm shift brought about by foundation models. By integrating attention mechanisms and leveraging large-scale pre-training, Transformers and foundation models enable more robust, interpretable, and scalable anomaly detection solutions. This work provides a comprehensive review of state-of-the-art techniques, their strengths, limitations, and emerging trends in leveraging these architectures for VAD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。