用图像分组构建双锚点,提升零样本异常检测稳定性。
Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection

- 引入自顶向下分组机制生成图像锚点,增强视觉语义表征。
- 在8个工业和6个医疗数据集上实现领先性能,显著降低对提示依赖。
- 适合需要高鲁棒性异常检测的工业与医疗场景应用。
零样本异常检测(ZSAD)旨在识别未见领域中的异常,这一设定在工业和医疗应用中尤为重要,因领域偏移普遍存在。然而,多数基于CLIP的ZSAD方法仅依赖文本模态锚定语义,导致性能高度依赖提示设计且视觉定位能力弱。为此,本文提出双锚点框架,通过自顶向下分组机制构建层次化图像锚点,逐步聚合局部到全局图像特征,形成正常与异常组令牌作为图像锚点,并作为门控信号输入组门控令牌重构器,以增强全局表征。重构后的图像锚点与文本提示融合生成动态状态提示。通过联合强化视觉与文本语义,该框架稳定了图文对齐,减少提示依赖,在8个工业与6个医疗基准上实现强泛化性能。
原文摘要 · Abstract (English)
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen domains, a setting that is particularly critical for industrial and medical applications where domain shifts are prevalent. However, most CLIP-based ZSAD methods anchor semantics solely on the text modality, making performance highly sensitive to prompt design and leading to weak visual grounding. To mitigate these limitations, we propose a Dual-Anchor framework that complements conventional text anchors with hierarchical image anchors constructed via a top-down grouping mechanism. This mechanism progressively aggregates local-to-global image features to form normal and abnormal group tokens, which serve as image anchors and act as gating signals in a Group-Gated Token Refiner to enhance the global representation. The refined image anchors are then fused with text prompts to construct dynamic state prompts. By jointly reinforcing visual and textual semantics, our framework stabilizes image-text alignment, reduces prompt dependency, and achieves strong generalization across 8 industrial and 6 medical benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。