arXiv:2412.06510cs.CVcs.AI2024-12被引 9

通过跨模态语义特征控制异常合成,提升真实感与可控性。

AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis

  • 采用文本-图像参考对提取跨模态语义特征作为引导。
  • 在多个数据集上生成的异常样本达到当前最优真实感表现。
  • 适合需要高精度异常数据增强的工业质检场景使用。

异常合成是提升异常检测性能的重要数据增强手段。现有基于大规模预训练的文本到图像异常合成方法主要依赖文本信息或粗粒度视觉特征进行生成,难以捕捉真实异常的细粒度视觉模式,限制了生成结果的真实性和泛化能力。为此,本文提出AnomalyControl框架,通过学习跨模态语义特征作为引导信号,编码来自文本-图像参考提示的通用异常线索,以提升合成异常样本的真实感。具体地,AnomalyControl采用灵活的非匹配提示对(即参考文本-图像提示与目标文本提示),设计跨模态语义建模(CSM)模块,从文本和视觉描述中提取跨模态语义特征;进一步提出异常语义增强注意力(ASEA)机制,使CSM聚焦于异常的特定视觉模式,增强生成特征的真实感与上下文相关性;最后,将跨模态语义特征作为先验,设计语义引导适配器(SGA),编码有效的引导信号以实现充分且可控的合成过程。大量实验表明,AnomalyControl在异常合成任务中优于现有方法,并在下游任务中展现出更优性能。

原文摘要 · Abstract (English)

Anomaly synthesis is a crucial approach to augment abnormal data for advancing anomaly inspection. Based on the knowledge from the large-scale pre-training, existing text-to-image anomaly synthesis methods predominantly focus on textual information or coarse-aligned visual features to guide the entire generation process. However, these methods often lack sufficient descriptors to capture the complicated characteristics of realistic anomalies (e.g., the fine-grained visual pattern of anomalies), limiting the realism and generalization of the generation process. To this end, we propose a novel anomaly synthesis framework called AnomalyControl to learn cross-modal semantic features as guidance signals, which could encode the generalized anomaly cues from text-image reference prompts and improve the realism of synthesized abnormal samples. Specifically, AnomalyControl adopts a flexible and non-matching prompt pair (i.e., a text-image reference prompt and a targeted text prompt), where a Cross-modal Semantic Modeling (CSM) module is designed to extract cross-modal semantic features from the textual and visual descriptors. Then, an Anomaly-Semantic Enhanced Attention (ASEA) mechanism is formulated to allow CSM to focus on the specific visual patterns of the anomaly, thus enhancing the realism and contextual relevance of the generated anomaly features. Treating cross-modal semantic features as the prior, a Semantic Guided Adapter (SGA) is designed to encode effective guidance signals for the adequate and controllable synthesis process. Extensive experiments indicate that AnomalyControl can achieve state-of-the-art results in anomaly synthesis compared with existing methods while exhibiting superior performance for downstream tasks.

异常合成跨模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。