通过分离前景背景特征,实现零样本细粒度异常检测
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
- 用多策略文本表示增强语义描述能力
- 通过多维分离与背景抑制提升定位精度
- 适合工业和医疗场景下的无监督异常识别
细粒度异常检测在工业和医疗应用中至关重要,但标注异常数据通常稀缺,导致零样本检测困难。尽管视觉语言模型如 CLIP 提供了有前景的解决方案,但仍面临前景-背景特征纠缠和文本语义粗略的问题。我们提出 FB-CLIP 框架,通过多策略文本表示和前景-背景分离来增强异常定位能力。在文本模态中,融合句尾特征、全局池化表示与注意力加权标记特征,以获取更丰富的语义线索。在视觉模态中,沿身份、语义和空间维度进行多视角软分离,并结合背景抑制,降低干扰并提升判别力。语义一致性正则化(SCR)将图像特征对齐至正常与异常的文本原型,抑制不确定匹配并扩大语义差距。实验表明,FB-CLIP 能有效区分复杂背景中的异常,实现在零样本设置下的准确细粒度异常检测与定位。
原文摘要 · Abstract (English)
Fine-grained anomaly detection is crucial in industrial and medical applications, but labeled anomalies are often scarce, making zero-shot detection challenging. While vision-language models like CLIP offer promising solutions, they struggle with foreground-background feature entanglement and coarse textual semantics. We propose FB-CLIP, a framework that enhances anomaly localization via multi-strategy textual representations and foreground-background separation. In the textual modality, it combines End-of-Text features, global-pooled representations, and attention-weighted token features for richer semantic cues. In the visual modality, multi-view soft separation along identity, semantic, and spatial dimensions, together with background suppression, reduces interference and improves discriminability. Semantic Consistency Regularization (SCR) aligns image features with normal and abnormal textual prototypes, suppressing uncertain matches and enlarging semantic gaps. Experiments show that FB-CLIP effectively distinguishes anomalies from complex backgrounds, achieving accurate fine-grained anomaly detection and localization under zero-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。