提出双分支结构分离全局与局部语义,提升视频时间定位精度。
Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding
- 分路处理句级与短语级语义,显式解耦全局与局部信息。
- 在QVHighlights和Charades-STA上刷新基准,提升关键指标。
- 适合需要细粒度视频-语言对齐的研究者或开发者。
视频时间定位(VTG)旨在从长且未剪辑的视频中定位与自然语言查询匹配的时间片段,通常包含两个子任务:时刻检索(MR)和亮点检测(HD)。尽管近年来基于CLIP和InternVideo2等预训练视觉-语言模型取得进展,但现有方法在跨模态注意力中对所有文本标记一视同仁,忽视其不同的语义角色。通过受控实验验证,当前模型过度依赖[EOS]驱动的全局语义,而未能有效利用词级信号,限制了细粒度时间对齐能力。为此,本文提出DualGround,一种双分支架构:将[EOS]标记通过句级路径处理,同时将词标记聚类为短语级单元以实现局部定位。方法引入两种机制:(1) 基于词角色的跨模态交互策略,结构化解耦地对齐视频特征与句级、短语级语义;(2) 联合建模范式,在增强句级对齐的同时,通过结构化短语感知上下文提升细粒度时间定位。该设计使模型能捕捉粗粒度与局部语义,实现更富表达力和上下文感知的视频定位。DualGround在QVHighlights和Charades-STA两个基准上均达到最先进性能,证明了解耦语义建模在视频-语言对齐中的有效性。
原文摘要 · Abstract (English)
Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: Moment Retrieval (MR) and Highlight Detection (HD). While recent advances have been progressed by powerful pretrained vision-language models such as CLIP and InternVideo2, existing approaches commonly treat all text tokens uniformly during crossmodal attention, disregarding their distinct semantic roles. To validate the limitations of this approach, we conduct controlled experiments demonstrating that VTG models overly rely on [EOS]-driven global semantics while failing to effectively utilize word-level signals, which limits their ability to achieve fine-grained temporal alignment. Motivated by this limitation, we propose DualGround, a dual-branch architecture that explicitly separates global and local semantics by routing the [EOS] token through a sentence-level path and clustering word tokens into phrase-level units for localized grounding. Our method introduces (1) tokenrole- aware cross modal interaction strategies that align video features with sentence-level and phrase-level semantics in a structurally disentangled manner, and (2) a joint modeling framework that not only improves global sentence-level alignment but also enhances finegrained temporal grounding by leveraging structured phrase-aware context. This design allows the model to capture both coarse and localized semantics, enabling more expressive and context-aware video grounding. DualGround achieves state-of-the-art performance on both Moment Retrieval and Highlight Detection tasks across QVHighlights and Charades- STA benchmarks, demonstrating the effectiveness of disentangled semantic modeling in video-language alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。