arXiv:2510.26464cs.CV2025-10被引 4

通过细粒度文本描述提升图像异常定位精度

Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection

  • 用多层级细粒度描述自动构建文本标签
  • 在MVTec-AD和VisA上实现更优定位效果
  • 适合关注少样本异常检测的科研人员

少样本异常检测(FSAD)方法利用少量正常样本识别异常区域。现有方法依赖预训练视觉语言模型(VLMs)通过图文特征相似性识别潜在异常区域,但因缺乏详细文本描述,仅能使用图像级描述匹配每个视觉补丁,导致语义与补丁级异常不匹配,定位性能受限。为此,我们提出多层级细粒度语义描述(MFSC),通过自动化流程为现有异常检测数据集构建多层级、细粒度文本描述。基于MFSC,提出新框架FineGrainedAD,包含两个组件:多层级可学习提示(MLLP)和多层级语义对齐(MLSA)。MLLP通过自动替换与拼接机制引入细粒度语义至多层级可学习提示;MLSA设计区域聚合策略与多层级对齐训练,使可学习提示更好对齐对应视觉区域。实验表明,所提FineGrainedAD在少样本设置下于MVTec-AD和VisA数据集上均取得更优整体性能。

原文摘要 · Abstract (English)

Few-shot anomaly detection (FSAD) methods identify anomalous regions with few known normal samples. Most existing methods rely on the generalization ability of pre-trained vision-language models (VLMs) to recognize potentially anomalous regions through feature similarity between text descriptions and images. However, due to the lack of detailed textual descriptions, these methods can only pre-define image-level descriptions to match each visual patch token to identify potential anomalous regions, which leads to the semantic misalignment between image descriptions and patch-level visual anomalies, achieving sub-optimal localization performance. To address the above issues, we propose the Multi-Level Fine-Grained Semantic Caption (MFSC) to provide multi-level and fine-grained textual descriptions for existing anomaly detection datasets with automatic construction pipeline. Based on the MFSC, we propose a novel framework named FineGrainedAD to improve anomaly localization performance, which consists of two components: Multi-Level Learnable Prompt (MLLP) and Multi-Level Semantic Alignment (MLSA). MLLP introduces fine-grained semantics into multi-level learnable prompts through automatic replacement and concatenation mechanism, while MLSA designs region aggregation strategy and multi-level alignment training to facilitate learnable prompts better align with corresponding visual regions. Experiments demonstrate that the proposed FineGrainedAD achieves superior overall performance in few-shot settings on MVTec-AD and VisA datasets.

异常检测视觉语言模型少样本学习细粒度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。