arXiv:2503.17651cs.CV2025-03

用视频与语言的协同时间一致性,提升仅标注一帧的视频定位精度。

Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization

  • 设计帧级与段级时序一致性模块,增强视频语义与文本对齐
  • 跨路径一致性引导使两个学习路径相互强化,提升定位准确率
  • 适合低标注成本场景下的视频-语言定位任务

自然语言视频定位(NLVL)旨在根据语言描述定位视频中的目标时刻。近期提出的点监督范式仅需目标时刻内一个标注帧,平衡了定位精度与标注成本。然而,因缺乏完整边界标注,视频内容与语言描述对齐困难,影响定位准确性。为此,本文提出协同时间一致性学习框架(COTEL),通过显著性检测与定位任务的协同,强化视频-语言对齐。设计帧级与段级时序一致性学习模块,建模帧显著性与句子-时刻对之间的语义一致性;引入跨一致性引导机制(FCG与SCG),使两条学习路径相互增强;并提出分层对比对齐损失(HCAL),全面优化对齐效果。在两个基准数据集上的实验表明,本方法优于现有最先进方法。代码将公开。

原文摘要 · Abstract (English)

Natural language video localization (NLVL) is a crucial task in video understanding that aims to localize the target moment in videos specified by a given language description. Recently, a point-supervised paradigm has been presented to address this task, requiring only a single annotated frame within the target moment rather than complete temporal boundaries. Compared with the fully-supervised paradigm, it offers a balance between localization accuracy and annotation cost. However, due to the absence of complete annotation, it is challenging to align the video content with language descriptions, consequently hindering accurate moment prediction. To address this problem, we propose a new COllaborative Temporal consistEncy Learning (COTEL) framework that leverages the synergy between saliency detection and moment localization to strengthen the video-language alignment. Specifically, we first design a frame- and a segment-level Temporal Consistency Learning (TCL) module that models semantic alignment across frame saliencies and sentence-moment pairs. Then, we design a cross-consistency guidance scheme, including a Frame-level Consistency Guidance (FCG) and a Segment-level Consistency Guidance (SCG), that enables the two temporal consistency learning paths to reinforce each other mutually. Further, we introduce a Hierarchical Contrastive Alignment Loss (HCAL) to comprehensively align the video and text query. Extensive experiments on two benchmarks demonstrate that our method performs favorably against SoTA approaches. We will release all the source codes.

视频定位点监督时序对齐多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。