构建首个聚焦视觉干扰下的逻辑异常检测数据集,助力工业质检突破外观变化干扰。
VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction
- 基于文本描述的对比学习框架,分离逻辑属性与低层视觉特征。
- 涵盖50个单类任务、10,395张图像,覆盖10种制造场景与5种拍摄条件。
- 适合关注工业视觉检测、逻辑推理与鲁棒性建模的研究者。
工业检测中的逻辑异常检测因视觉外观变化(如背景杂乱、光照偏移、模糊)而面临挑战,这些因素常使以视觉为中心的检测器偏离规则级违规识别。现有基准很少提供逻辑状态固定但干扰因素变化的可控设置。为此,我们提出VID-AD,一个面向视觉诱导干扰下的图像级逻辑异常检测数据集。该数据集包含10个制造场景和5种采集条件,共50个单类任务,10,395张图像。每个场景由数量、长度、类型、位置、关系五类逻辑约束中选出两个定义,异常包括单约束及组合违规。我们进一步提出一种基于语言的异常检测框架,仅依赖从正常图像生成的文本描述。通过正样本文本与基于矛盾的负样本文本进行对比学习,方法学习出捕捉逻辑属性而非低层特征的嵌入表示。大量实验表明,该方法在不同设置下均优于基线模型。数据集已开源:https://github.com/nkthiroto/VID-AD。
原文摘要 · Abstract (English)
Logical anomaly detection in industrial inspection remains challenging due to variations in visual appearance (e.g., background clutter, illumination shift, and blur), which often distract vision-centric detectors from identifying rule-level violations. However, existing benchmarks rarely provide controlled settings where logical states are fixed while such nuisance factors vary. To address this gap, we introduce VID-AD, a dataset for logical anomaly detection under vision-induced distraction. It comprises 10 manufacturing scenarios and five capture conditions, totaling 50 one-class tasks and 10,395 images. Each scenario is defined by two logical constraints selected from quantity, length, type, placement, and relation, with anomalies including both single-constraint and combined violations. We further propose a language-based anomaly detection framework that relies solely on text descriptions generated from normal images. Using contrastive learning with positive texts and contradiction-based negative texts synthesized from these descriptions, our method learns embeddings that capture logical attributes rather than low-level features. Extensive experiments demonstrate consistent improvements over baselines across the evaluated settings. The dataset is available at: https://github.com/nkthiroto/VID-AD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。