通过删句增强上下文一致性,提升视频段落定位的半监督性能
Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding
- 删去句子生成强扰动输入,引导学生模型学习教师信号
- 利用原始与增强视图预测一致度生成伪标签,提升训练可靠性
- 在有限标注下显著优于现有方法,适合弱监督视频理解场景
半监督视频段落定位(SSVPG)旨在仅用少量时间标注,从无剪辑视频中定位一段话中的多个句子。现有方法主要依赖教师-学生一致性学习和视频级对比损失,但忽略了通过扰动查询上下文来生成强监督信号的重要性。本文提出一种新的上下文一致性学习(CCL)框架,统一了一致性正则化与伪标签范式以增强半监督学习。具体地,首先进行教师-学生学习:学生模型接收移除部分句子的强增强样本,强制学习来自教师模型的强监督信号;随后基于生成的伪标签进行模型重训练,利用原始与增强视图预测间的互认同作为标签置信度。大量实验表明,CCL显著超越现有方法。
原文摘要 · Abstract (English)
Semi-Supervised Video Paragraph Grounding (SSVPG) aims to localize multiple sentences in a paragraph from an untrimmed video with limited temporal annotations. Existing methods focus on teacher-student consistency learning and video-level contrastive loss, but they overlook the importance of perturbing query contexts to generate strong supervisory signals. In this work, we propose a novel Context Consistency Learning (CCL) framework that unifies the paradigms of consistency regularization and pseudo-labeling to enhance semi-supervised learning. Specifically, we first conduct teacher-student learning where the student model takes as inputs strongly-augmented samples with sentences removed and is enforced to learn from the adequately strong supervisory signals from the teacher model. Afterward, we conduct model retraining based on the generated pseudo labels, where the mutual agreement between the original and augmented views' predictions is utilized as the label confidence. Extensive experiments show that CCL outperforms existing methods by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。