arXiv:2608.00986cs.CV2026-08

从费舍尔信息角度分析多模态异常检测中的融合偏差并提出解决方法

Understanding and Overcoming Cross-modal Fusion Bias in Multimodal Anomaly Detection From A Fisher Information Perspective

论文配图:Understanding and Overcoming Cross-modal Fusion Bias in Multimodal Anomaly Detection From A Fisher Information Perspective
图 1 · 摘自论文原文
  • 基于费舍尔信息矩阵分析跨模态融合偏差影响
  • 在MVTec 3D-AD和Eyecandies上实现单类/多类/少样本场景提升
  • 适合关注多模态学习偏差与鲁棒性优化的研究者

当前多模态异常检测(MAD)主要依赖增强跨模态融合,尤其是通过融合RGB与深度数据以获得更丰富的异常表征。然而,对跨模态融合偏差这一多模态学习中常见挑战在MAD中的作用关注不足。本文首次通过费舍尔信息矩阵分析该偏差的影响,并据此提出UCFB框架——一种简单有效的即插即用方案,通过费舍尔信息引导的动态校准调整模态特定正则化权重,并结合典型相似性分析改善模态间交互。在MVTec 3D-AD和Eyecandies数据集上的大量实验表明,UCFB在单类、多类及少样本设置下均取得一致性能提升。

原文摘要 · Abstract (English)

Current advancements in Multimodal Anomaly Detection (MAD) are largely driven by enhancing multimodal fusion, particularly through the integration of RGB and Depth data for richer anomaly representation. However, less attention was devoted to analyzing the role of cross-modal fusion bias, a well-known challenge in multimodal learning, in MAD. This gap motivates a key question: can we overcome this bias to break the performance bottleneck of current work? In this paper, we first analyze the impact of cross-modal fusion bias in MAD via the Fisher Information Matrix. Then, grounded in these findings, we propose UCFB, a simple yet effective plug-and-play framework designed to mitigate cross-modal fusion bias in MAD. It achieves this by jointly employing Fisher-information-guided dynamic calibration to adjust modality-specific regularization weights and canonical similarity analysis to improve inter-modal interactions. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets demonstrate that UCFB achieves consistent improvements in single-class, multi-class, and few-shot settings.

多模态异常检测融合偏差费舍尔信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。