提升弱监督视频定位精度,通过分阶段学习让模型逐步掌握复杂查询。
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
- 分阶段训练+上下文感知,逐步提升对复杂查询的理解能力。
- 在VUHM、TGIF-QA等数据集上显著优于基线方法,尤其在长视频中表现突出。
- 适合研究视频理解、多模态学习的开发者和研究人员参考。
本文聚焦弱监督时空视频定位(WSTVG),该任务旨在无边界框标注情况下,根据文本查询定位特定主体的时空位置。尽管当前先进的目标检测模型具备强大的零样本能力,但其在时间预测不一致、复杂查询理解不足及难场景适应性差方面仍存在明显局限。为此,本文提出一种新方法CoSPaL(上下文自步学习),包含三个核心组件:(1) 管道短语定位(TPG),通过将文本查询与时空管状体关联实现定位;(2) 上下文指代定位(CRG),利用上下文信息增强对象识别时序一致性;(3) 自步场景理解(SPS),采用渐进式训练策略,从粗粒度到细粒度逐步提升模型应对复杂场景的能力。实验表明,CoSPaL在VUHM、TGIF-QA等多个基准上显著超越现有方法,尤其在长视频和复杂查询场景下优势明显。
原文摘要 · Abstract (English)
In this work, we focus on Weakly Supervised Spatio-Temporal Video Grounding (WSTVG). It is a multimodal task aimed at localizing specific subjects spatio-temporally based on textual queries without bounding box supervision. Motivated by recent advancements in multi-modal foundation models for grounding tasks, we first explore the potential of state-of-the-art object detection models for WSTVG. Despite their robust zero-shot capabilities, our adaptation reveals significant limitations, including inconsistent temporal predictions, inadequate understanding of complex queries, and challenges in adapting to difficult scenarios. We propose CoSPaL (Contextual Self-Paced Learning), a novel approach which is designed to overcome these limitations. CoSPaL integrates three core components: (1) Tubelet Phrase Grounding (TPG), which introduces spatio-temporal prediction by linking textual queries to tubelets; (2) Contextual Referral Grounding (CRG), which improves comprehension of complex queries by extracting contextual information to refine object identification over time; and (3) Self-Paced Scene Understanding (SPS), a training paradigm that progressively increases task difficulty, enabling the model to adapt to complex scenarios by transitioning from coarse to fine-grained understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。