通过渐进式学习提升弱监督时空定位的准确率
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
- 分阶段训练:先学动作片段,再处理复杂场景
- 在三个数据集上达到最优,最高提升3.0%
- 适合做视频理解与多模态定位的研究者
本文研究弱监督时空视频定位(WSTVG),即仅用文本查询定位视频中目标的时空位置,无需边界框标注。受视觉语言基础模型启发,我们探索其零样本定位能力,但发现直接适配缺乏关键的时空定位性能。为此提出管段引用定位(TRG)方法,将文本与视频管段关联以实现时空预测。然而,TRG在组合动作理解和密集场景中表现不佳。为此,我们提出新型渐进学习框架STPro,包含两个核心模块:(1) 子动作时间课程学习(SA-TCL),逐步构建组合动作理解能力;(2) 拥挤引导空间课程学习(CG-SCL),通过空间难度递增适应复杂场景。STPro在三个基准数据集上取得当前最优结果,在VidSTG-Declarative上提升1.0%,在HCSTVG-v1上提升3.0%。
原文摘要 · Abstract (English)
In this work we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries and no bounding box supervision. Inspired by recent advances in vision-language foundation models, we investigate their utility for WSTVG, leveraging their zero-shot grounding capabilities. However, we find that a simple adaptation lacks essential spatio-temporal grounding abilities. To bridge this gap, we introduce Tubelet Referral Grounding (TRG), which connects textual queries to tubelets to enable spatio-temporal predictions. Despite its promise, TRG struggles with compositional action understanding and dense scene scenarios. To address these limitations, we propose STPro, a novel progressive learning framework with two key modules: (1) Sub-Action Temporal Curriculum Learning (SA-TCL), which incrementally builds compositional action understanding, and (2) Congestion-Guided Spatial Curriculum Learning (CG-SCL), which adapts the model to complex scenes by spatially increasing task difficulty. STPro achieves state-of-the-art results on three benchmark datasets, with improvements of 1.0% on VidSTG-Declarative and 3.0% on HCSTVG-v1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。