arXiv:2608.24937cs.LGcs.AI2026-08中稿 · publication in IEE…综述

系统梳理多模态异常检测的主流方法与未来方向

Multi-Modal Anomaly Detection: A Survey

论文配图:Multi-Modal Anomaly Detection: A Survey
图 1 · 摘自论文原文
  • 按正常性与异常性假设分类,提出两类核心检测范式
  • 揭示五项内在挑战,整合跨领域多模态异常检测方法
  • 适合关注工业、安防等领域异常检测的研究者参考

多模态异常检测(MMAD)从异构数据源中识别罕见异常事件,广泛应用于工业质检和网络安全等关键场景。现有研究分散于不同领域与模态组合,传统综述常按架构分类,未能反映异常定义与分离机制的本质差异。本文从假设驱动视角出发,形式化问题,提炼出五个核心挑战,并将已有工作归纳为两大互补范式:正常性假设方法通过表征学习、跨模态对齐与知识增强建模常规模式;异常性假设方法则通过粗粒度、结构化和语义异常注入,优化决策边界。此外,我们探讨基础模型如何通过可扩展预训练、灵活跨模态迁移和新兴推理能力重塑MMAD。最后,整理代表性基准与评估协议,指出鲁棒、自适应与可解释性系统的关键开放问题与未来方向。

原文摘要 · Abstract (English)

Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-critical applications such as industrial inspection and cybersecurity. Yet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnormality is defined and separated in multi-modal settings. We survey MMAD from an assumption-driven perspective. We formalize the problem, identify five intrinsic characteristics underlying its core challenges, and organize prior work into two complementary paradigms. The first, normality-assumption methods, models regularity via representation learning, cross-modal alignment, and knowledge enhancement. The second, anomaly-assumption methods, sharpens decision boundaries through coarse-grained, structural, and semantic anomaly injection. We also investigate how foundation models are reshaping MMAD through scalable pretraining, flexible cross-modal transfer, and emerging reasoning capabilities. Finally, we compile representative benchmarks and evaluation protocols across domains and highlight open problems and future directions for robust, adaptive, and interpretable MMAD systems.

异常检测多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。