通过投票机制提升动作边界定位精度,解决标注模糊问题。
Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation

- 构建层级投票网络,用全局先验增强局部过渡区域特征
- 在三个数据集上显著提升边界定位准确率,改善伪标签质量
- 适合需要高精度动作分割的视频理解任务
时间戳监督的动作分割旨在仅以每段动作随机一帧标注的情况下,对未剪辑视频进行动作分割与分类。精确地从时间戳标注中定位动作边界至关重要,因为这能生成帧级伪标签,并应用成熟的全监督训练方法。然而,现有方法在动作转换区域因特征不明显而面临固有不确定性,导致边界估计不准确,从而降低生成伪标签的稳定性和可靠性,最终影响分割模型性能。本文提出边界投票网络,通过分层传播视频级全局先验知识到局部动作转换区域,缓解特征模糊问题。通过在整个视频中生成关键动作表示作为‘投票’,并聚焦于动作转换区域,所有投票协同增强该区域特征并精炼边界定位。大量实验表明,该方法在GTEA、50Salads和Breakfast数据集上均有效。
原文摘要 · Abstract (English)
Timestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a random frame annotated per action. Precisely localizing action boundaries from timestamp annotations is crucial for this setting, as it enables generating framewise pseudo-labels and applying the well-explored fully-supervised training. However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions. This imprecise boundary estimation significantly reduces the stability and reliability of the generated pseudo-labels in ambiguous action-transiting regions, consequently resulting in performance deterioration of the trained segmentation models. In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. By generating key action representations as votes throughout the video and targeting action-transiting regions, all votes collaboratively contribute to action-transiting feature enhancement and boundary localization refinement. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。