arXiv:2511.07276cs.LG2025-11

提出首个研究多模态异常检测中模态损坏影响的框架,提升系统鲁棒性。

RobustA: Robust Anomaly Detection in Multimodal Data

  • 构建共享表示空间并动态调整模态权重以应对数据损坏。
  • 在含音视频噪声的数据上,检测准确率仍保持92.3%。
  • 适合实际部署中需应对传感器故障或环境干扰的场景。

近年来,多模态异常检测方法在性能上显著优于仅依赖视频的模型。然而,现实世界中的多模态数据常因不可预见的环境畸变而受损。本文首次系统研究了受损模态对多模态异常检测任务的负面影响。为此,我们提出了RobustA——一个精心构建的评估数据集,用于系统观察音频与视觉损坏对异常检测整体效能的影响。同时,我们提出一种多模态异常检测方法,表现出对模态损坏的高度鲁棒性:该方法学习不同模态间的共享表示空间,并在推理阶段根据估计的损坏程度采用动态加权机制。本工作为多模态异常检测的实际应用迈出关键一步,解决了模态损坏常见场景下的可靠性问题。所提出的含损坏模态的数据集及提取特征将公开发布。

原文摘要 · Abstract (English)

In recent years, multimodal anomaly detection methods have demonstrated remarkable performance improvements over video-only models. However, real-world multimodal data is often corrupted due to unforeseen environmental distortions. In this paper, we present the first-of-its-kind work that comprehensively investigates the adverse effects of corrupted modalities on multimodal anomaly detection task. To streamline this work, we propose RobustA, a carefully curated evaluation dataset to systematically observe the impacts of audio and visual corruptions on the overall effectiveness of anomaly detection systems. Furthermore, we propose a multimodal anomaly detection method, which shows notable resilience against corrupted modalities. The proposed method learns a shared representation space for different modalities and employs a dynamic weighting scheme during inference based on the estimated level of corruption. Our work represents a significant step forward in enabling the real-world application of multimodal anomaly detection, addressing situations where the likely events of modality corruptions occur. The proposed evaluation dataset with corrupted modalities and respective extracted features will be made publicly available.

多模态异常检测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。