生成逼真异常图像与标注掩码,解决结构不一致和特征纠缠问题。
Double Helix Diffusion for Cross-Domain Anomaly Image Generation
- 双螺旋架构分离图像与掩码特征,独立增强
- 生成图像与掩码结构一致,真实感更强
- 支持文本提示控制,适合工业异常检测应用
视觉异常检测在制造业中至关重要,但受限于真实异常样本稀缺,难以训练鲁棒的检测器。合成数据生成是数据增强的有效策略,但现有方法存在两大瓶颈:1)生成的异常与正常背景在结构上不一致;2)合成图像与其对应标注掩码间存在不良特征纠缠,影响输出的感知真实性。本文提出双螺旋扩散模型(DH-Diff),一种新型跨域生成框架,可同时生成高保真异常图像及其像素级标注掩码,针对性解决上述问题。该模型采用受双螺旋启发的独特架构,循环遍历特征分离、连接与融合模块。具体地,领域解耦注意力机制通过独立增强图像与标注特征来缓解特征纠缠;语义评分图对齐模块则通过协同整合异常前景,确保结构真实性。DH-Diff支持文本提示与可选图形引导,具备灵活控制能力。大量实验表明,其在多样性与真实性上显著优于当前最先进方法,进而大幅提升下游异常检测性能。
原文摘要 · Abstract (English)
Visual anomaly inspection is critical in manufacturing, yet hampered by the scarcity of real anomaly samples for training robust detectors. Synthetic data generation presents a viable strategy for data augmentation; however, current methods remain constrained by two principal limitations: 1) the generation of anomalies that are structurally inconsistent with the normal background, and 2) the presence of undesirable feature entanglement between synthesized images and their corresponding annotation masks, which undermines the perceptual realism of the output. This paper introduces Double Helix Diffusion (DH-Diff), a novel cross-domain generative framework designed to simultaneously synthesize high-fidelity anomaly images and their pixel-level annotation masks, explicitly addressing these challenges. DH-Diff employs a unique architecture inspired by a double helix, cycling through distinct modules for feature separation, connection, and merging. Specifically, a domain-decoupled attention mechanism mitigates feature entanglement by enhancing image and annotation features independently, and meanwhile a semantic score map alignment module ensures structural authenticity by coherently integrating anomaly foregrounds. DH-Diff offers flexible control via text prompts and optional graphical guidance. Extensive experiments demonstrate that DH-Diff significantly outperforms state-of-the-art methods in diversity and authenticity, leading to significant improvements in downstream anomaly detection performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。