arXiv:2504.14860cs.CVcs.AI2025-04CVPR被引 11

用伪标签提升弱监督动作定位性能,缩小与全监督差距

Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormer

  • 双分支架构融合片段与提案级先验生成高质量伪标签
  • 在THUMOS14和ActivityNet1.3上达到当前最优性能
  • 适合关注弱监督视频理解的算法研究者

弱监督时间动作定位(WTAL)虽取得显著进展,但仍因缺乏时间标注导致性能与全监督方法存在差距。现有方法使用伪标签训练时,仍面临三大挑战:生成高质量伪标签、充分利用不同先验信息、在噪声标签下优化训练策略。为此,本文提出PseudoFormer,一种新型双分支框架,旨在弥合弱监督与全监督时间动作定位之间的差距。首先引入RickerFusion,将所有预测动作提案映射到全局共享空间,生成更优伪标签;随后,利用弱分支提供的片段级与提案级标签,结合不同先验信息,训练全分支中的回归模型;最后,采用不确定性掩码与迭代精炼机制应对噪声伪标签。PseudoFormer在两个常用基准数据集THUMOS14与ActivityNet1.3上均取得当前最优结果。大量消融实验验证了各模块的有效性。

原文摘要 · Abstract (English)

Weakly-supervised Temporal Action Localization (WTAL) has achieved notable success but still suffers from a lack of temporal annotations, leading to a performance and framework gap compared with fully-supervised methods. While recent approaches employ pseudo labels for training, three key challenges: generating high-quality pseudo labels, making full use of different priors, and optimizing training methods with noisy labels remain unresolved. Due to these perspectives, we propose PseudoFormer, a novel two-branch framework that bridges the gap between weakly and fully-supervised Temporal Action Localization (TAL). We first introduce RickerFusion, which maps all predicted action proposals to a global shared space to generate pseudo labels with better quality. Subsequently, we leverage both snippet-level and proposal-level labels with different priors from the weak branch to train the regression-based model in the full branch. Finally, the uncertainty mask and iterative refinement mechanism are applied for training with noisy pseudo labels. PseudoFormer achieves state-of-the-art WTAL results on the two commonly used benchmarks, THUMOS14 and ActivityNet1.3. Besides, extensive ablation studies demonstrate the contribution of each component of our method.

动作定位弱监督伪标签视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。