arXiv:2409.09953cs.CV2024-09中稿 · MIPR 2024被引 12

提出新模型检测视频中异常动作,兼顾外观与运动信息。

Uncertainty-Guided Appearance-Motion Association Network for Out-of-Distribution Action Detection

  • 分离提取外观与运动特征,构建时空图捕捉物体交互。
  • 在两个数据集上显著优于现有方法,提升动作识别准确率。
  • 适合需要鲁棒视频理解的场景,如安防监控、自动驾驶。

分布外(OOD)检测旨在识别并拒绝语义发生偏移的测试样本,防止在分布内(ID)数据集上训练的模型产生不可靠预测。现有工作仅在图像数据集上提取外观特征,难以处理包含大量运动信息的动态多媒体场景。因此,本文聚焦更真实且更具挑战性的任务:分布外动作检测(ODAD)。给定一段未修剪视频,ODAD需先分类出ID动作并识别出OOD动作,再定位这些动作。为此,本文提出一种新型不确定性引导的外观-运动关联网络(UAAN),同时利用外观特征与运动上下文,推理时空中的物体间交互关系。首先,设计独立的外观与运动分支,分别提取面向外观和运动的物体表示;每个分支中构建时空图以推理外观引导与运动驱动的物体间交互。随后,设计外观-运动注意力模块融合两者特征以完成最终动作检测。在两个具有挑战性的数据集上的实验结果表明,UAAN显著超越现有最先进方法,证明其有效性。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection targets to detect and reject test samples with semantic shifts, to prevent models trained on in-distribution (ID) dataset from producing unreliable predictions. Existing works only extract the appearance features on image datasets, and cannot handle dynamic multimedia scenarios with much motion information. Therefore, we target a more realistic and challenging OOD detection task: OOD action detection (ODAD). Given an untrimmed video, ODAD first classifies the ID actions and recognizes the OOD actions, and then localizes ID and OOD actions. To this end, in this paper, we propose a novel Uncertainty-Guided Appearance-Motion Association Network (UAAN), which explores both appearance features and motion contexts to reason spatial-temporal inter-object interaction for ODAD.Firstly, we design separate appearance and motion branches to extract corresponding appearance-oriented and motion-aspect object representations. In each branch, we construct a spatial-temporal graph to reason appearance-guided and motion-driven inter-object interaction. Then, we design an appearance-motion attention module to fuse the appearance and motion features for final action detection. Experimental results on two challenging datasets show that UAAN beats state-of-the-art methods by a significant margin, illustrating its effectiveness.

动作检测视频理解分布外检测多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。