arXiv:2503.07253cs.CV2025-03被引 7

用视觉语言模型和扩散模型生成逼真多样的工业缺陷图像

AnomalyPainter: Vision-Language-Diffusion Synergy for Zero-Shot Realistic and Diverse Industrial Anomaly Synthesis

  • 结合大模型与纹理库,自动生成缺陷描述并匹配多样纹理
  • 在多个数据集上实现更高真实感与多样性,优于现有方法
  • 适合需要零样本生成工业缺陷数据的研究与工业质检场景

现有异常合成方法虽有进展,但在真实感与多样性之间仍存在权衡。为此,我们提出 AnomalyPainter,一个零样本框架,通过融合视觉语言大模型(VLLM)、潜在扩散模型(LDM)及新构建的纹理库 Tex-9K 实现突破。Tex-9K 包含75类、8,792个专业纹理资产,专为多样化异常合成设计。利用 VLLM 的通用知识,为每类工业物体生成合理缺陷文本描述,并从 Tex-9K 中匹配相关纹理。这些纹理通过 ControlNet 指导 LDM 在正常图像上进行绘制。此外,引入 Texture-Aware Latent Init 以稳定基于自然图像训练的 ControlNet 在工业图像上的表现。大量实验表明,AnomalyPainter 在真实感、多样性和泛化能力上均优于现有方法,下游任务性能显著提升。

原文摘要 · Abstract (English)

While existing anomaly synthesis methods have made remarkable progress, achieving both realism and diversity in synthesis remains a major obstacle. To address this, we propose AnomalyPainter, a zero-shot framework that breaks the diversity-realism trade-off dilemma through synergizing Vision Language Large Model (VLLM), Latent Diffusion Model (LDM), and our newly introduced texture library Tex-9K. Tex-9K is a professional texture library containing 75 categories and 8,792 texture assets crafted for diverse anomaly synthesis. Leveraging VLLM's general knowledge, reasonable anomaly text descriptions are generated for each industrial object and matched with relevant diverse textures from Tex-9K. These textures then guide the LDM via ControlNet to paint on normal images. Furthermore, we introduce Texture-Aware Latent Init to stabilize the natural-image-trained ControlNet for industrial images. Extensive experiments show that AnomalyPainter outperforms existing methods in realism, diversity, and generalization, achieving superior downstream performance.

异常生成扩散模型零样本工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。