arXiv:2504.05700cs.CV2025-04CVPR被引 2

利用人体姿态信息提升弱监督动作分割效果

Pose-Aware Weakly-Supervised Action Segmentation

  • 训练时引入姿态知识,推理时不使用,实现知识蒸馏
  • 在多个数据集上优于现有最先进方法,提升显著
  • 兼容多种模型结构,适合长视频动作分割任务

理解人类行为是视觉智能的重要目标。其中一大挑战在于准确标注动作片段需要大量人工成本。为此,本文提出一种弱监督框架,在长视频动作分割中仅需极少标注。该方法在训练阶段融合人体姿态信息,但推理时不依赖姿态,从而将与各动作相关的姿态知识有效提炼。设计了一种基于姿态的对比损失,增强对动作边界的区分能力。在多个代表性数据集上的大量实验表明,该方法在在线与离线设置下均超越现有最先进水平。同时,框架对不同分割主干网络和姿态提取器具有良好的适应性。

原文摘要 · Abstract (English)

Understanding human behavior is an important problem in the pursuit of visual intelligence. A challenge in this endeavor is the extensive and costly effort required to accurately label action segments. To address this issue, we consider learning methods that demand minimal supervision for segmentation of human actions in long instructional videos. Specifically, we introduce a weakly-supervised framework that uniquely incorporates pose knowledge during training while omitting its use during inference, thereby distilling pose knowledge pertinent to each action component. We propose a pose-inspired contrastive loss as a part of the whole weakly-supervised framework which is trained to distinguish action boundaries more effectively. Our approach, validated through extensive experiments on representative datasets, outperforms previous state-of-the-art (SOTA) in segmenting long instructional videos under both online and offline settings. Additionally, we demonstrate the framework's adaptability to various segmentation backbones and pose extractors across different datasets.

动作分割弱监督姿态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。