用弱监督方法实现跨视频动作片段的精准分割。
2by2: Weakly-Supervised Learning for Global Action Segmentation
- 基于视频对的三元学习,利用活动标签构建动作表征。
- 在Breakfast和YouTube Instructions数据集上超越现有方法。
- 适合关注弱监督动作识别与视频理解的研究者。
本文提出一种简单而有效的方法,解决尚不充分研究的全局动作分割任务,旨在将不同活动视频中表示相同动作的帧进行分组。由于视频间动作顺序不一致,该任务更具挑战性。我们利用活动标签,以弱监督方式学习适用于全局动作分割的动作表征。为此,提出一种视频对的三元学习方法,确保视频内动作区分性以及视频间、活动间的动作关联性。主干网络采用基于稀疏Transformer的孪生网络,输入视频对并判断是否属于同一活动。该方法在Breakfast和YouTube Instructions两个基准数据集上验证,性能优于现有最先进方法。
原文摘要 · Abstract (English)
This paper presents a simple yet effective approach for the poorly investigated task of global action segmentation, aiming at grouping frames capturing the same action across videos of different activities. Unlike the case of videos depicting all the same activity, the temporal order of actions is not roughly shared among all videos, making the task even more challenging. We propose to use activity labels to learn, in a weakly-supervised fashion, action representations suitable for global action segmentation. For this purpose, we introduce a triadic learning approach for video pairs, to ensure intra-video action discrimination, as well as inter-video and inter-activity action association. For the backbone architecture, we use a Siamese network based on sparse transformers that takes as input video pairs and determine whether they belong to the same activity. The proposed approach is validated on two challenging benchmark datasets: Breakfast and YouTube Instructions, outperforming state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。