针对微小缺陷难标注难题,提出融合生成与语义的主动学习框架
Hard to See, Hard to Label: Generative and Symbolic Acquisition for Subtle Visual Phenomena

- 用扩散模型评估图像重建差异和去噪变异,识别视觉模糊样本
- 在三个数据集上显著提升罕见缺陷检测效率,标签节省超30%
- 适合工业质检场景,尤其对低频、难辨缺陷有强捕捉能力
细微视觉异常如细裂纹、亚毫米空洞和低对比度夹杂物,结构异常但视觉模糊,难以标注且易被忽略。基于判别不确定性或特征多样性的传统采集方法常过度选择主导模式,忽视稀疏但关键的数据区域。在工业缺陷检测中,此类异常往往低频且难以区分。为此,我们提出GSAL,一种结合扩散难度信号与分层语义覆盖先验的主动学习框架。扩散组件通过重建误差与去噪变异性评分,优先选择视觉异常或模糊样本;但仅靠扩散会反复聚焦于主导语义模式中的难例。因此,语义组件构建三级概念图,促进对未充分覆盖语义区域的探索,并提供可解释的采样理由。通过平衡视觉难度与语义覆盖,GSAL在专有薄膜缺陷数据集、Pascal VOC和MS COCO上均实现标签效率提升,优于基于不确定性和多样性等基线方法,显著增强对细微稀有目标的检出能力。
原文摘要 · Abstract (English)
Subtle visual anomalies such as hairline cracks, sub-millimeter voids, and low-contrast inclusions are structurally atypical yet visually ambiguous, making them both difficult to annotate and easy to overlook during active learning. Standard acquisition heuristics based on discriminative uncertainty or feature diversity often overselect dominant patterns while underexploring sparse yet important regions of the data space. This failure mode is especially severe in industrial defect inspection, where anomalies may be both low-prevalence and difficult to distinguish from surrounding structure. To resolve this, we propose GSAL, an active learning framework for object detection that combines a diffusion-based difficulty signal with a hierarchical semantic coverage prior. The diffusion component scores images and proposals using reconstruction discrepancy and denoising variability, prioritizing visually atypical or ambiguous examples. However, diffusion alone does not prevent acquisition from repeatedly favoring hard samples within dominant semantic modes. The semantic component therefore organizes candidate samples in a three-level concept graph and promotes coverage of underrepresented semantic regions while providing interpretable acquisition rationales. By balancing visual difficulty with semantic coverage, GSAL improves retrieval of subtle and rare targets that are often missed by uncertainty-only selection. Experiments on a proprietary thin-film defect, Pascal VOC and MS COCO dataset show consistent gains in label efficiency and rare-class retrieval over uncertainty-, diversity-, and hybrid-based baselines
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。