针对工业缺陷检测,融合视觉与文本信息提升少样本异常识别效果
VTFusion: A Vision-Text Multimodal Fusion Network for Few-Shot Anomaly Detection
- 自适应提取图像和文本特征,学习工业场景专用表示
- 2样本下在MVTec AD和VisA上达到96.8%和86.2%的图像级AUROC
- 适用于真实工业场景,对汽车塑料件缺陷检测达93.5% AUPRO
少样本异常检测(FSAD)已成为利用有限正常样本识别异常的关键范式。现有方法虽引入文本语义补充视觉信息,但多依赖自然场景预训练特征,忽视工业检测所需的细粒度领域语义。同时,主流融合策略常采用简单拼接,无法解决视觉与文本模态间的语义错位问题,易受跨模态干扰。为此,本文提出VTFusion,一种面向FSAD的视觉-文本多模态融合框架。其核心包括:1)为图像与文本模态设计自适应特征提取器,学习任务特定表示,弥合预训练模型与工业数据间的领域差距,并通过生成多样合成异常增强特征可区分性;2)构建专用多模态预测融合模块,包含促进跨模态深度交互的融合块及在多模态引导下生成精细像素级异常图的分割网络。VTFusion显著提升FSAD性能,在MVTec AD和VisA数据集2样本场景下分别取得96.8%和86.2%的图像级AUROC。此外,在本文引入的真实工业汽车塑料件数据集上,实现93.5%的AUPRO,充分验证其在严苛工业场景中的实际应用价值。
原文摘要 · Abstract (English)
Few-Shot Anomaly Detection (FSAD) has emerged as a critical paradigm for identifying irregularities using scarce normal references. While recent methods have integrated textual semantics to complement visual data, they predominantly rely on features pre-trained on natural scenes, thereby neglecting the granular, domain-specific semantics essential for industrial inspection. Furthermore, prevalent fusion strategies often resort to superficial concatenation, failing to address the inherent semantic misalignment between visual and textual modalities, which compromises robustness against cross-modal interference. To bridge these gaps, this study proposes VTFusion, a vision-text multimodal fusion framework tailored for FSAD. The framework rests on two core designs. First, adaptive feature extractors for both image and text modalities are introduced to learn task-specific representations, bridging the domain gap between pre-trained models and industrial data; this is further augmented by generating diverse synthetic anomalies to enhance feature discriminability. Second, a dedicated multimodal prediction fusion module is developed, comprising a fusion block that facilitates rich cross-modal information exchange and a segmentation network that generates refined pixel-level anomaly maps under multimodal guidance. VTFusion significantly advances FSAD performance, achieving image-level AUROCs of 96.8% and 86.2% in the 2-shot scenario on the MVTec AD and VisA datasets, respectively. Furthermore, VTFusion achieves an AUPRO of 93.5% on a real-world dataset of industrial automotive plastic parts introduced in this paper, further demonstrating its practical applicability in demanding industrial scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。