用文本控制风格迁移,零样本生成局部异常图像。
AnoStyler: Text-Driven Localized Anomaly Generation via Lightweight Style Transfer
- 将异常生成转化为文本引导的轻量级风格迁移任务。
- 在MVTec-AD和VisA上生成图像质量高、多样性好。
- 仅需一张正常图+文本提示,适合快速构建异常数据集。
异常生成被广泛用于解决真实数据中异常图像稀缺的问题。然而,现有方法通常存在至少以下一种局限:(1)生成异常视觉真实性不足;(2)依赖大量真实图像;(3)使用内存密集型重型模型架构。为克服这些限制,我们提出AnoStyler,一种轻量且高效的方法,将零样本异常生成建模为文本引导的风格迁移。给定单张正常图像及其类别标签和期望缺陷类型,通过通用的类别无关流程生成表示异常区域的异常掩码以及描述正常与异常状态的双类文本提示。采用基于CLIP损失函数训练的轻量级U-Net模型,将正常图像风格化为视觉逼真的异常图像,其中异常由异常掩码定位,并与文本提示语义对齐。在MVTec-AD和VisA数据集上的大量实验表明,AnoStyler在生成高质量、多样化的异常图像方面优于现有方法。此外,使用生成的异常图像可有效提升异常检测性能。
原文摘要 · Abstract (English)
Anomaly generation has been widely explored to address the scarcity of anomaly images in real-world data. However, existing methods typically suffer from at least one of the following limitations, hindering their practical deployment: (1) lack of visual realism in generated anomalies; (2) dependence on large amounts of real images; and (3) use of memory-intensive, heavyweight model architectures. To overcome these limitations, we propose AnoStyler, a lightweight yet effective method that frames zero-shot anomaly generation as text-guided style transfer. Given a single normal image along with its category label and expected defect type, an anomaly mask indicating the localized anomaly regions and two-class text prompts representing the normal and anomaly states are generated using generalizable category-agnostic procedures. A lightweight U-Net model trained with CLIP-based loss functions is used to stylize the normal image into a visually realistic anomaly image, where anomalies are localized by the anomaly mask and semantically aligned with the text prompts. Extensive experiments on the MVTec-AD and VisA datasets show that AnoStyler outperforms existing anomaly generation methods in generating high-quality and diverse anomaly images. Furthermore, using these generated anomalies helps enhance anomaly detection performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。