arXiv:2608.13973cs.CV2026-08

用辅助模态条件化增强RGB图像异常检测,不破坏原有语义路径。

Rethinking Auxiliary Modalities in Multimodal Zero-shot Anomaly Detection: From Semantic Fusion to Conditional Modulation

论文配图:Rethinking Auxiliary Modalities in Multimodal Zero-shot Anomaly Detection: From Semantic Fusion to Conditional Modulation
图 1 · 摘自论文原文
  • 以辅助模态作为条件信号,动态优化RGB特征而非融合语义空间。
  • 在MVTec 3D-AD和Eyecandies上提升多个主流RGB检测器性能,达最新水平。
  • 轻量级模块支持即插即用,适合已有零样本异常检测系统升级。

基于基础模型的近期方法通过视觉语言预训练赋予RGB图像强大的零样本异常检测能力。然而,仅依赖RGB观测仍难以捕捉由几何形变、深度变化或细微表面变化主导的异常。辅助模态可提供互补结构信息,但现有多模态方法通常将它们直接融合到共享语义空间,可能干扰由RGB基础模型建立的文本对齐异常语义,且常需特定模态架构。为此,本文提出一种即插即用的辅助条件增强框架。不重构联合多模态异常语义空间,该框架保留原始RGB图像-文本匹配路径,利用辅助观测作为条件信号来精细化RGB特征,使辅助模态无缝增强现有基于RGB的零样本异常检测器。具体地,一个轻量级元学习模块以全局RGB与辅助表示为输入,生成样本自适应的低秩残差更新,决定如何调整RGB特征;进一步从初始RGB异常响应和辅助可靠性构建不确定性感知的空间调制,决定局部残差更新的增强或抑制。这种从全局到局部的条件调制实现选择性多模态增强,同时保持原始RGB异常语义。在MVTec 3D-AD和Eyecandies上的大量实验表明,本框架持续提升多个主流基于RGB的零样本异常检测器,在多模态零样本异常检测中达到领先性能。

原文摘要 · Abstract (English)

Recent foundation model-based methods have endowed RGB images with strong zero-shot anomaly detection (ZSAD) through vision-language pretraining. However, RGB observations alone remain limited in perceiving anomalies dominated by geometric deformation, depth variation, or subtle surface changes. Auxiliary modalities can provide complementary structural information, but existing multimodal methods typically fuse them directly into a shared semantic space, which may disturb the text-aligned anomaly semantics established by RGB foundation models and often requires modality-specific architectures. To address this issue, we propose a plug-and-play auxiliary-conditioned enhancement framework for zero-shot anomaly detection. Instead of reconstructing a joint multimodal anomaly semantic space, our framework preserves the original RGB image-text anomaly matching pathway and uses auxiliary observations as conditional signals for RGB feature refinement, allowing auxiliary modalities to seamlessly enhance existing RGB-based zero-shot anomaly detectors. Specifically, a lightweight meta-learning module takes global RGB and auxiliary representations as input and generates sample-adaptive low-rank residual updates to determine how RGB features should be refined. We further construct uncertainty-aware spatial modulation from the initial RGB anomaly response and auxiliary reliability, which determines where local residual updates are strengthened or suppressed. This global-to-local conditional modulation enables selective multimodal enhancement while preserving the original RGB anomaly semantics. Extensive experiments on MVTec 3D-AD and Eyecandies demonstrate that our framework consistently improves multiple popular RGB-based zero-shot anomaly detectors, achieving state-of-the-art performance for multimodal zero-shot anomaly detection.

异常检测多模态零样本条件增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。