arXiv:2412.19418cs.CVcs.AI2024-12被引 10

通过融合不确定度与多头注意力,提升弱监督动作定位的准确率

Generalized Uncertainty-Based Evidential Fusion with Hybrid Multi-Head Attention for Weak-Supervised Temporal Action Localization

  • 设计混合多头注意力模块,优化视频特征表达
  • 基于不确定度的证据融合使定位精度提升1.8% (mAP)
  • 适合处理背景噪声干扰的弱监督动作定位任务

弱监督时间动作定位(WS-TAL)旨在仅使用视频级标签定位完整动作实例并进行分类。现有方法面临动作-背景模糊性挑战,主要由聚合过程中的背景噪声及动作内部差异引起。本文提出混合多头注意力(HMHA)模块与广义不确定度驱动的证据融合(GUEF)模块。HMHA通过过滤冗余信息并调整特征分布,增强RGB与光流特征表现;GUEF通过融合片段级证据,自适应消除背景噪声干扰,改进不确定性度量并选择更优前景特征,使模型聚焦于完整动作实例。在THUMOS14数据集上的实验表明,该方法优于当前最优方法,实现1.8%的mAP提升。

原文摘要 · Abstract (English)

Weakly supervised temporal action localization (WS-TAL) is a task of targeting at localizing complete action instances and categorizing them with video-level labels. Action-background ambiguity, primarily caused by background noise resulting from aggregation and intra-action variation, is a significant challenge for existing WS-TAL methods. In this paper, we introduce a hybrid multi-head attention (HMHA) module and generalized uncertainty-based evidential fusion (GUEF) module to address the problem. The proposed HMHA effectively enhances RGB and optical flow features by filtering redundant information and adjusting their feature distribution to better align with the WS-TAL task. Additionally, the proposed GUEF adaptively eliminates the interference of background noise by fusing snippet-level evidences to refine uncertainty measurement and select superior foreground feature information, which enables the model to concentrate on integral action instances to achieve better action localization and classification performance. Experimental results conducted on the THUMOS14 dataset demonstrate that our method outperforms state-of-the-art methods. Our code is available in \url{https://github.com/heyuanpengpku/GUEF/tree/main}.

动作定位弱监督注意力机制不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。