arXiv:2507.07579cs.CVcs.AI2025-07

用视觉大模型+多任务学习,实现少样本跨域缺陷检测新突破

NexViTAD: Few-shot Unsupervised Cross-Domain Defect Detection via Vision Foundation Models and Multi-Task Learning

论文配图:NexViTAD: Few-shot Unsupervised Cross-Domain Defect Detection via Vision Foundation Models and Multi-Task Learning
图 1 · 摘自论文原文
  • 融合Hiera与DINO-v2特征的分层适配器构建鲁棒表征
  • 通过瓶颈约束和跳跃连接实现跨域知识迁移,性能超主流模型
  • 支持多源域并行处理,适合工业场景少样本缺陷检测

本文提出一种基于视觉基础模型的少样本跨域异常检测框架NexViTAD,通过创新的共享子空间投影机制和多任务学习模块,有效应对工业异常检测中的域偏移问题。核心创新包括:(1) 分层适配器模块,自适应融合Hiera与DINO-v2预训练模型的互补特征,构建更鲁棒的特征表示;(2) 共享子空间投影策略,通过瓶颈维度约束与跳跃连接机制实现高效跨域知识迁移;(3) 多任务解码器架构,支持多源域同时处理,显著提升模型泛化能力;(4) 基于Sinkhorn-K-means聚类的异常分数推理方法,结合高斯滤波与自适应阈值处理,实现像素级精准定位。在MVTec AD数据集上,NexViTAD在目标域取得97.5% AUC、70.4% AP和95.2% PRO的领先性能,超越现有模型,标志着跨域缺陷检测的重要进展。

原文摘要 · Abstract (English)

This paper presents a novel few-shot cross-domain anomaly detection framework, Nexus Vision Transformer for Anomaly Detection (NexViTAD), based on vision foundation models, which effectively addresses domain-shift challenges in industrial anomaly detection through innovative shared subspace projection mechanisms and multi-task learning (MTL) module. The main innovations include: (1) a hierarchical adapter module that adaptively fuses complementary features from Hiera and DINO-v2 pre-trained models, constructing more robust feature representations; (2) a shared subspace projection strategy that enables effective cross-domain knowledge transfer through bottleneck dimension constraints and skip connection mechanisms; (3) a MTL Decoder architecture supports simultaneous processing of multiple source domains, significantly enhancing model generalization capabilities; (4) an anomaly score inference method based on Sinkhorn-K-means clustering, combined with Gaussian filtering and adaptive threshold processing for precise pixel level. Valuated on the MVTec AD dataset, NexViTAD delivers state-of-the-art performance with an AUC of 97.5%, AP of 70.4%, and PRO of 95.2% in the target domains, surpassing other recent models, marking a transformative advance in cross-domain defect detection.

缺陷检测跨域学习视觉大模型少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。