arXiv:2602.05718cs.CV2026-02被引 1

通过自监督任务提升点标注动作定位的时序一致性理解能力

Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization

  • 设计三项自监督时序理解任务,挖掘点标注中的时序信息
  • 在四个数据集上显著优于现有方法,平均mAP提升3.2%以上
  • 适合关注弱监督视频动作定位与时序建模的研究者

点标注时序动作定位(PTAL)采用轻量级帧标注范式(即每段动作仅标注一帧)训练模型,在未剪辑视频中定位动作实例。现有方法通常仅使用点标注进行片段级分类,缺乏对动作内帧间时序关系的显式建模。然而,理解帧间的时序关系对于把握动作定义至关重要,有助于定位完整动作区间。为此,本文提出一种多任务学习框架,充分利用点标注信息以增强模型对动作时序一致性的理解能力。具体设计三项自监督时序理解任务:(i) 动作补全,(ii) 动作顺序理解,(iii) 动作规律性理解。这些任务共同促进模型对跨视频动作时序一致性的建模。据我们所知,这是首次在点标注动作定位中显式探索时序一致性的尝试。在四个基准数据集上的大量实验表明,该方法显著优于多个先进方法。

原文摘要 · Abstract (English)

Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per action instance) to train a model to effectively locate action instances within untrimmed videos. Most existing approaches design the task head of models with only a point-supervised snippet-level classification, without explicit modeling of understanding temporal relationships among frames of an action. However, understanding the temporal relationships of frames is crucial because it can help a model understand how an action is defined and therefore benefits localizing the full frames of an action. To this end, in this paper, we design a multi-task learning framework that fully utilizes point supervision to boost the model's temporal understanding capability for action localization. Specifically, we design three self-supervised temporal understanding tasks: (i) Action Completion, (ii) Action Order Understanding, and (iii) Action Regularity Understanding. These tasks help a model understand the temporal consistency of actions across videos. To the best of our knowledge, this is the first attempt to explicitly explore temporal consistency for point supervision action localization. Extensive experimental results on four benchmark datasets demonstrate the effectiveness of the proposed method compared to several state-of-the-art approaches.

动作定位弱监督时序建模自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。