arXiv:2505.17555cs.HCcs.CV2025-05中稿 · CHI'25被引 3

用户拖拽关联人体与物体,自动标注视频动作时序,大幅降低人工标注成本。

ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization

  • 通过拖拽节点并连接定义动作关键事件及其空间关系。
  • 在无标注视频上生成大量带标签数据,支持半监督训练模型。
  • 适合视频理解研究者与需要快速构建标注体系的开发者。

时序动作定位(Temporal Action Localization, TAL)旨在检测视频中动作的起止时间。然而,训练TAL模型需要大量人工标注数据。数据编程是一种通过人工定义标注函数高效生成训练标签的方法,但在处理时序视频中的复杂动作时面临挑战。本文提出ProTAL,一种面向TAL的拖拽-链接视频编程框架。用户可通过拖拽代表人体部位和物体的节点,并通过连线约束其空间关系(如方向、距离等),定义关键事件。这些定义被用于为大规模无标注视频生成动作标签,再采用半监督方法训练TAL模型。我们通过使用场景演示和用户研究验证了ProTAL的有效性,为视频编程框架设计提供了新思路。

原文摘要 · Abstract (English)

Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to create training labels with a series of human-defined labeling functions. However, its application in TAL faces difficulties of defining complex actions in the context of temporal video frames. In this paper, we propose ProTAL, a drag-and-link video programming framework for TAL. ProTAL enables users to define \textbf{key events} by dragging nodes representing body parts and objects and linking them to constrain the relations (direction, distance, etc.). These definitions are used to generate action labels for large-scale unlabelled videos. A semi-supervised method is then employed to train TAL models with such labels. We demonstrate the effectiveness of ProTAL through a usage scenario and a user study, providing insights into designing video programming framework.

动作定位视频编程半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。