用可学习的语义锚点提升大模型零样本异常分割精度
AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
- 引入三类可学习语义锚点,将抽象异常概念转为可定位的视觉实体
- 在6个工业与医疗数据集上实现零样本异常分割新纪录
- 适合需要高精度异常检测的工业质检与医学影像分析场景
大型多模态模型(LMMs)具备强大的任务泛化能力,为零样本视觉异常分割(ZSAS)带来新机遇。然而现有方法仍存在根本局限:异常概念抽象且依赖上下文,缺乏稳定的视觉原型;高层语义嵌入与像素级空间特征对齐薄弱,影响精确定位。为此,本文提出AG-VAS框架,通过引入三个可学习的语义锚点-[SEG]、[NOR]和[ANO],建立统一的锚点引导分割范式。其中,[SEG]作为绝对语义锚点,将抽象异常语义转化为具象的空间视觉实体(如孔洞或划痕);[NOR]和[ANO]作为相对锚点,建模跨类别正常与异常模式的上下文对比。为进一步增强跨模态对齐,设计了语义-像素对齐模块(SPAM)和锚点引导掩码解码器(AGMD),实现精准异常定位。此外,构建了Anomaly-Instruct20K大规模指令数据集,以结构化方式组织外观、形状和空间属性等异常知识,支持语义锚点的有效学习与融合。在六个工业与医疗基准上的大量实验表明,AG-VAS在零样本设置下持续达到最先进性能。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) exhibit strong task generalization capabilities, offering new opportunities for zero-shot visual anomaly segmentation (ZSAS). However, existing LMM-based segmentation approaches still face fundamental limitations: anomaly concepts are inherently abstract and context-dependent, lacking stable visual prototypes, and the weak alignment between high-level semantic embeddings and pixel-level spatial features hinders precise anomaly localization. To address these challenges, we present AG-VAS (Anchor-Guided Visual Anomaly Segmentation), a new framework that expands the LMM vocabulary with three learnable semantic anchor tokens-[SEG], [NOR], and [ANO], establishing a unified anchor-guided segmentation paradigm. Specifically, [SEG] serves as an absolute semantic anchor that translates abstract anomaly semantics into explicit, spatially grounded visual entities (e.g., holes or scratches), while [NOR] and [ANO] act as relative anchors that model the contextual contrast between normal and abnormal patterns across categories. To further enhance cross-modal alignment, we introduce a Semantic-Pixel Alignment Module (SPAM) that aligns language-level semantic embeddings with high-resolution visual features, along with an Anchor-Guided Mask Decoder (AGMD) that performs anchor-conditioned mask prediction for precise anomaly localization. In addition, we curate Anomaly-Instruct20K, a large-scale instruction dataset that organizes anomaly knowledge into structured descriptions of appearance, shape, and spatial attributes, facilitating effective learning and integration of the proposed semantic anchors. Extensive experiments on six industrial and medical benchmarks demonstrate that AG-VAS achieves consistent state-of-the-art performance in the zero-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。