arXiv:2604.08301cs.CV2026-04

用空间语义控制生成异常图像,解决少样本下异常合成不准的问题。

GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis

论文配图:GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis
图 1 · 摘自论文原文
  • 引入空间条件模块,利用像素级语义图精准控制异常位置。
  • 在MVTec AD和VisA上达到最新水平,异常检测与分割性能领先。
  • 适合工业质检中真实异常数据稀缺的场景,可快速适配新缺陷类型。

工业质量控制中的视觉异常检测性能常受限于真实异常样本的稀缺。为此,异常合成技术被用来扩充训练集并提升下游检测效果。然而,现有方法或因图像修复导致融合不佳,或无法生成精确掩码。为此,我们提出GroundingAnomaly——一种新颖的少样本异常图像生成框架。该框架引入空间条件模块,利用像素级语义图实现对合成异常的空间精准控制;同时设计门控自注意力模块,通过门控注意力层将条件标记注入冻结的U-Net,有效保留预训练先验并确保稳定少样本适应。在MVTec AD和VisA数据集上的大量实验表明,GroundingAnomaly能生成高质量异常图像,并在异常检测、分割及实例级检测等下游任务中达到当前最优表现。

原文摘要 · Abstract (English)

The performance of visual anomaly inspection in industrial quality control is often constrained by the scarcity of real anomalous samples. Consequently, anomaly synthesis techniques have been developed to enlarge training sets and enhance downstream inspection. However, existing methods either suffer from poor integration caused by inpainting or fail to provide accurate masks. To address these limitations, we propose GroundingAnomaly, a novel few-shot anomaly image generation framework. Our framework introduces a Spatial Conditioning Module that leverages per-pixel semantic maps to enable precise spatial control over the synthesized anomalies. Furthermore, a Gated Self-Attention Module is designed to inject conditioning tokens into a frozen U-Net via gated attention layers. This carefully preserves pretrained priors while ensuring stable few-shot adaptation. Extensive evaluations on the MVTec AD and VisA datasets demonstrate that GroundingAnomaly generates high-quality anomalies and achieves state-of-the-art performance across multiple downstream tasks, including anomaly detection, segmentation, and instance-level detection.

异常生成扩散模型少样本学习工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。