arXiv:2603.18481cs.CVcs.LG2026-03被引 3

提出T-QPM框架,提升视觉语言模型在动态环境中的异常检测能力

T-QPM: Enabling Temporal Out-Of-Distribution Detection and Domain Generalization for Vision-Language Models in Open-World

  • 引入时序四重模式匹配,结合跨模态一致性增强判断边界
  • 通过轻量级融合权重应对时间分布漂移,提升稳定性
  • 适合开放世界中需持续适应新数据的视觉语言系统

开放世界学习中,模型需应对不断变化的数据分布,异常检测仍是关键挑战。现有视觉语言模型如CLIP虽能通过双模式匹配(DPM)实现多模态异常检测,但普遍存在两大缺陷:(1)依赖固定融合规则,假设环境静态,难以应对时间漂移;(2)对协变量分布偏移鲁棒性差。本文提出两步式新框架,将双模式扩展为时序四重模式匹配(T-QPM)。首先,通过将异常图像与文本描述配对,建立异类信号间的跨模态一致性模式,利用联合图文推理优化决策边界;其次,针对时间分布漂移,学习轻量级融合权重,以最优结合语义匹配与视觉典型性。为保证稳定性,引入基于平均阈值置信度(ATC)的显式正则化,防止分布演化导致性能下降。在时间划分基准上的实验表明,该方法显著优于静态基线,在非平稳环境中提供稳健、时序一致的多模态异常检测框架。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection remains a critical challenge in open-world learning, where models must adapt to evolving data distributions. While recent vision-language models (VLMS) like CLIP enable multimodal OOD detection through Dual-Pattern Matching (DPM), existing methods typically suffer from two major shortcomings: (1) They rely on fixed fusion rules and assume static environments, failing under temporal drift; and (2) they lack robustness against covariate shifted inputs. In this paper, we propose a novel two-step framework to enhance OOD detection and covariate distribution shift robustness in dynamic settings. We extend the dual-pattern regime into Temporal Quadruple-Pattern Matching (T-QPM). First, by pairing OOD images with text descriptions, we introduce cross-modal consistency patterns between ID and OOD signals, refining the decision boundary through joint image-text reasoning. Second, we address temporal distribution shifts by learning lightweight fusion weights to optimally combine semantic matching and visual typicality. To ensure stability, we enforce explicit regularization based on Average Thresholded Confidence (ATC), preventing performance degradation as distributions evolve. Experiments on temporally partitioned benchmarks demonstrate that our approach significantly outperforms static baselines, offering a robust, temporally-consistent framework for multimodal OOD detection in non-stationary environments.

视觉语言模型异常检测开放世界时序鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。