用异构超图建模视觉与语义关系,提升少样本异常检测性能
H2VLR: Heterogeneous Hypergraph Vision-Language Reasoning for Few-Shot Anomaly Detection

- 构建视觉区域与语义概念的统一超图,实现高阶关系推理
- 在工业与医疗数据集上达到当前最优效果,显著优于传统匹配方法
- 适合需要少样本学习的工业质检和医学影像分析场景
异常检测是工业质检和医学影像中的经典视觉任务,但常面临数据稀缺问题。为此,少样本异常检测(FSAD)受到广泛关注。近年来,视觉语言模型(VLM)被引入以提升性能,但现有方法大多仅依赖成对特征匹配,忽略了结构依赖与全局一致性。为此,本文提出异构超图视觉语言推理(H2VLR)框架,将FSAD重构为视觉-语义关系的高阶推理问题,通过统一超图联合建模视觉区域与语义概念。实验表明,H2VLR在代表性工业与医学基准上均达当前最优(SOTA)性能。代码将在论文接受后发布。
原文摘要 · Abstract (English)
As a classic vision task, anomaly detection has been widely applied in industrial inspection and medical imaging. In this task, data scarcity is often a frequently-faced issue. To solve it, the few-shot anomaly detection (FSAD) scheme is attracting increasing attention. In recent years, beyond traditional visual paradigm, Vision-Language Model (VLM) has been extensively explored to boost this field. However, in currently-existing VLM-based FSAD schemes, almost all perform anomaly inference only by pairwise feature matching, ignoring structural dependencies and global consistency. To further redound to FSAD via VLM, we propose a Heterogeneous Hypergraph Vision-Language Reasoning (H2VLR) framework. It reformulates the FSAD as a high-order inference problem of visual-semantic relations, by jointly modeling visual regions and semantic concepts in a unified hypergraph. Experimental comparisons verify the effectiveness and advantages of H2VLR. It could often achieve state-of-the-art (SOTA) performance on representative industrial and medical benchmarks. Our code will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。