arXiv:2606.24404cs.CV2026-06中稿 · ECCV

针对多模态动作识别,提出新方法提升异常检测能力。

Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition

论文配图:Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition
图 1 · 摘自论文原文
  • 利用多模态与单模态预测间的关联关系设计后处理检测器。
  • 在多个数据集上平均性能超越现有最佳方法。
  • 适合需要鲁棒多模态识别的工业应用开发者。

将额外模态引入动作识别模型可显著提升其在多种场景下的表现。然而,如何利用这些信息增强模型鲁棒性仍不明确,尤其是在多模态分布外(OOD)检测方面。现有方法虽在训练时考虑了OOD检测,但推理阶段仍使用为单模态设计的现成检测器,导致信息浪费。我们发现多模态与单模态预测间存在重要关联,据此提出一种专为多模态场景设计的后处理检测器。该方法结合特征空间得分与多模态逻辑值进行归一化,实现对多模态流形外样本的检测。所提混合检测器兼容现有训练策略,并在多模态OOD基准的多个主流数据集上实现平均性能超越当前最优水平。结果表明,推理阶段显式考虑不同模态对多模态OOD检测至关重要。

原文摘要 · Abstract (English)

The incorporation of additional modalities into action recognition models increases their performance across a wide range of settings. However, how this additional information can contribute to making the models more robust remains underexplored, particularly for the case of multi-modal out-of-distribution (OOD) detection. While methods exist that regularize the multi-modal training process with OOD detection in mind, they still apply off-the-shelf OOD detectors designed for the uni-modal case during inference, discarding important information. Based on an interesting relationship we find between the multi-modal and uni-modal predictions, we propose to use this signal to build a post-hoc detector explicitly designed for the multi-modal scenario. We combine this new source of information with a feature-space score, which detects off-manifold samples in the multi-modal space, and normalize them by the multi-modal logits. In doing so, the proposed hybrid detector is compatible with existing training-time approaches and consistently improves performance. Experiments on a wide range of established datasets from the MultiOOD benchmark show that, on average, our approach outperforms the state of the art. Our results show the importance of explicitly considering the different modalities at inference time for multi-modal OOD detection.

多模态异常检测动作识别深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。