通过多尺度时间信息提升弱监督动作定位精度
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
- 分两阶段利用多尺度时间信息生成高质量伪标签
- 在两个阶段间交替优化,显著提升定位准确率
- 适合关注弱监督视频理解的研究者和工程师
弱监督时间动作定位因训练时仅提供视频级标注而具有挑战性。本文提出一种两阶段方法,充分挖掘时间域的多分辨率信息,基于外观与运动流生成高质量帧级伪标签。第一阶段通过初始伪标签生成(ILG)模块,利用时间多分辨率一致性生成类激活序列(CAS),衡量每帧属于特定动作类的可能性。第二阶段提出渐进式时间伪标签精炼(PTLR)框架,采用原始时间尺度网络(Network-OTS)与降采样时间尺度网络(Network-RTS)作为双流结构,分别生成两类尺度下的CAS,并在两者间交替精炼伪标签。通过在伪标签层面交换多分辨率信息,相互增强,从而提升各流的预测性能。实验表明该方法在多个基准数据集上显著优于现有方法。
原文摘要 · Abstract (English)
Weakly supervised temporal action localization is a challenging task as only the video-level annotation is available during the training process. To address this problem, we propose a two-stage approach to fully exploit multi-resolution information in the temporal domain and generate high quality frame-level pseudo labels based on both appearance and motion streams. Specifically, in the first stage, we generate reliable initial frame-level pseudo labels, and in the second stage, we iteratively refine the pseudo labels and use a set of selected frames with highly confident pseudo labels to train neural networks and better predict action class scores at each frame. We fully exploit temporal information at multiple scales to improve temporal action localization performance. Specifically, in order to obtain reliable initial frame-level pseudo labels, in the first stage, we propose an Initial Label Generation (ILG) module, which leverages temporal multi-resolution consistency to generate high quality class activation sequences (CASs), which consist of a number of sequences with each sequence measuring how likely each video frame belongs to one specific action class. In the second stage, we propose a Progressive Temporal Label Refinement (PTLR) framework. In our PTLR framework, two networks called Network-OTS and Network-RTS, which are respectively used to generate CASs for the original temporal scale and the reduced temporal scales, are used as two streams (i.e., the OTS stream and the RTS stream) to refine the pseudo labels in turn. By this way, the multi-resolution information in the temporal domain is exchanged at the pseudo label level, and our work can help improve each stream (i.e., the OTS/RTS stream) by exploiting the refined pseudo labels from another stream (i.e., the RTS/OTS stream).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。