无需异常样本,用跨模态提示生成逼真异常图像。
Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation
- 通过视觉与文本联合编码,用提示词引导图像修复生成异常。
- 在12,987组数据上训练,生成异常更真实多样,提升检测准确率。
- 支持任意正常图像生成,适合做异常生成的通用基础模型。
我们提出Anomagic,一种零样本异常生成方法,无需任何异常样本即可生成语义一致的异常图像。通过跨模态提示编码机制融合视觉与文本信息,驱动基于图像修复的生成流程。后续对比精炼策略强化生成异常与掩码间的精准对齐,从而提升下游异常检测性能。为支持训练,我们构建了AnomVerse,一个包含12,987个异常-掩码-标题三元组的数据集,源自13个公开数据集,标题由多模态大模型基于结构化视觉提示和模板化文本提示自动生成。大量实验表明,基于AnomVerse训练的Anomagic可生成比现有方法更真实、更丰富的异常图像,显著提升下游异常检测效果。此外,Anomagic可基于用户定义提示,在任意正常类别图像上生成异常,建立了一个通用的异常生成基础模型。
原文摘要 · Abstract (English)
We propose Anomagic, a zero-shot anomaly generation method that produces semantically coherent anomalies without requiring any exemplar anomalies. By unifying both visual and textual cues through a crossmodal prompt encoding scheme, Anomagic leverages rich contextual information to steer an inpainting-based generation pipeline. A subsequent contrastive refinement strategy enforces precise alignment between synthesized anomalies and their masks, thereby bolstering downstream anomaly detection accuracy. To facilitate training, we introduce AnomVerse, a collection of 12,987 anomaly-mask-caption triplets assembled from 13 publicly available datasets, where captions are automatically generated by multimodal large language models using structured visual prompts and template-based textual hints. Extensive experiments demonstrate that Anomagic trained on AnomVerse can synthesize more realistic and varied anomalies than prior methods, yielding superior improvements in downstream anomaly detection. Furthermore, Anomagic can generate anomalies for any normal-category image using user-defined prompts, establishing a versatile foundation model for anomaly generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。