arXiv:2508.03539cs.CV2025-08AAAI被引 2

用文字精准生成局部缺陷图,提升异常检测真实性和效率。

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

  • 通过文本引导的自回归编辑,实现缺陷位置与语义的精确控制。
  • 在三个数据集上检测性能超越现有方法,合成速度提升5倍。
  • 适合需要高精度异常生成与检测的研究者使用。

尽管异常合成方法已取得显著进展,但现有的基于扩散模型和粗略修复的流程普遍存在结构缺陷,如微结构不连续、语义控制能力弱和生成效率低等问题。为此,我们提出ARAS,一种语言条件下的自回归异常合成方法,通过标记锚定的潜在空间编辑,将文本指定的局部缺陷精确注入正常图像中。借助硬门控自回归算子和无需训练的上下文保持掩码采样核,ARAS显著提升了缺陷的真实感,保留了细粒度材料纹理,并实现了对合成异常的连续语义控制。集成于我们的质量感知重加权异常检测(QARAD)框架中,我们进一步提出一种动态加权策略,通过双编码器模型计算图像-文本相似度,强调高质量合成样本。在三个基准数据集(MVTec AD、VisA、BTAD)上的大量实验表明,QARAD在图像级和像素级异常检测任务中均优于当前最优方法,准确率与鲁棒性更优,且相比基于扩散的方法合成速度提升5倍。完整代码与合成数据集将公开发布。

原文摘要 · Abstract (English)

Despite substantial progress in anomaly synthesis methods, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-structural discontinuities, limited semantic controllability, and inefficient generation. To overcome these limitations, we introduce ARAS, a language-conditioned, auto-regressive anomaly synthesis approach that precisely injects local, text-specified defects into normal images via token-anchored latent editing. Leveraging a hard-gated auto-regressive operator and a training-free, context-preserving masked sampling kernel, ARAS significantly enhances defect realism, preserves fine-grained material textures, and provides continuous semantic control over synthesized anomalies. Integrated within our Quality-Aware Re-weighted Anomaly Detection (QARAD) framework, we further propose a dynamic weighting strategy that emphasizes high-quality synthetic samples by computing an image-text similarity score with a dual-encoder model. Extensive experiments across three benchmark datasets-MVTec AD, VisA, and BTAD, demonstrate that our QARAD outperforms SOTA methods in both image- and pixel-level anomaly detection tasks, achieving improved accuracy, robustness, and a 5 times synthesis speedup compared to diffusion-based alternatives. Our complete code and synthesized dataset will be publicly available.

异常检测自回归文本生成图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。