arXiv:2412.21205cs.CVcs.AI2024-12AAAI被引 6

用少量自动采样帧标注,实现低成本高精度动作检测。

Action-Agnostic Point-Level Supervision for Temporal Action Detection

  • 自动采样视频帧,由人工仅标注动作类别。
  • 在多个数据集上表现优于或媲美传统标注方法。
  • 适合希望降低标注成本的研究者和工程师。

我们提出一种无动作依赖的点级监督(AAPL),用于时间动作检测,以实现仅需轻量标注的数据集下精确的动作实例检测。在该方案中,通过无监督方式采样少量视频帧,交由人工标注动作类别。与传统点级标注需人工搜索所有动作实例不同,本方法无需人工干预即可完成帧选择。我们还设计了相应的检测模型与学习方法,以高效利用这些AAPL标签。在THUMOS'14、FineAction、GTEA、BEOID和ActivityNet 1.3等多个数据集上的大量实验表明,该方法在标注成本与检测性能的权衡上,优于或媲美现有视频级与点级标注方法。

原文摘要 · Abstract (English)

We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who then label the frames with action categories. Unlike point-level supervision, which requires annotators to search for every action instance in an untrimmed video, frames to annotate are selected without human intervention in AAPL supervision. We also propose a detection model and learning method to effectively utilize the AAPL labels. Extensive experiments on the variety of datasets (THUMOS '14, FineAction, GTEA, BEOID, and ActivityNet 1.3) demonstrate that the proposed approach is competitive with or outperforms prior methods for video-level and point-level supervision in terms of the trade-off between the annotation cost and detection performance.

动作检测弱监督标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。