用部分标注提升视频时间定位精度,降低人工成本。
Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
- 设计对比统一框架,分隐式与显式两阶段精炼定位
- 在Charades-STA和ActivityNet上超越弱监督方法
- 适合标注资源有限但需高精度定位的场景
时间句定位旨在从无剪辑视频中检测自然语言查询描述的事件时间戳。现有全监督方法性能优异但依赖昂贵标注;弱监督使用廉价标签但效果不佳。为在低标注成本下获得高性能,本文提出一种中间部分监督设置——训练时仅提供短片段标签。为此设计对比统一框架,包含隐式-显式渐进定位两阶段:隐式阶段通过四重对比学习(事件-查询聚合、事件-背景分离、类内紧凑性、类间可分性)在细粒度上对齐事件-查询表征;显式阶段利用生成的高质量伪标签,训练全监督模型进行定位优化与去噪。在Charades-STA和ActivityNet Captions上的大量实验与详尽消融验证了部分监督的有效性及本方法的优越性能。
原文摘要 · Abstract (English)
Temporal sentence grounding aims to detect event timestamps described by the natural language query from given untrimmed videos. The existing fully-supervised setting achieves great results but requires expensive annotation costs; while the weakly-supervised setting adopts cheap labels but performs poorly. To pursue high performance with less annotation costs, this paper introduces an intermediate partially-supervised setting, i.e., only short-clip is available during training. To make full use of partial labels, we specially design one contrast-unity framework, with the two-stage goal of implicit-explicit progressive grounding. In the implicit stage, we align event-query representations at fine granularity using comprehensive quadruple contrastive learning: event-query gather, event-background separation, intra-cluster compactness and inter-cluster separability. Then, high-quality representations bring acceptable grounding pseudo-labels. In the explicit stage, to explicitly optimize grounding objectives, we train one fully-supervised model using obtained pseudo-labels for grounding refinement and denoising. Extensive experiments and thoroughly ablations on Charades-STA and ActivityNet Captions demonstrate the significance of partial supervision, as well as our superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。