arXiv:2605.26441cs.CVcs.AI2026-05ECCV被引 42

用博弈论新视角提升弱监督视频定位精度

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

论文配图:Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
图 1 · 摘自论文原文
  • 将帧与词建模为博弈参与者,动态评估跨模态匹配贡献
  • 在Charades-STA和ActivityNet上超越现有方法,无需依赖复杂候选框
  • 适合关注视频定位与跨模态对齐的 researchers

本文针对弱监督视频时间定位任务,提出一种基于博弈论的新范式。现有方法依赖对比学习与重建框架生成候选片段,但忽略了细粒度跨模态一致性及候选框质量对性能的影响。为此,本文首次将视频帧与查询词视为多变量合作博弈中的参与者,通过博弈交互量化帧-词组合在不同粒度下的协作趋势,从而评估所有可能的对应关系。最终,基于学习到的查询引导帧级得分进行定位,避免了对复杂候选框的依赖。在Charades-STA和ActivityNet Caption数据集上的实验表明,该方法显著优于现有基线。

原文摘要 · Abstract (English)

This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposal selection framework that utilizes contrastive learning and reconstruction paradigm for scoring the pre-defined moment proposals. Although they have achieved significant progress, we argue that their current frameworks have overlooked two indispensable issues: 1) Coarse-grained cross-modal learning: previous methods solely capture the global video-level alignment with the query, failing to model the detailed consistency between video frames and query words for accurately grounding the moment boundaries. 2) Complex moment proposals: their performance severely relies on the quality of proposals, which are also time-consuming and complicated for selection. To this end, in this paper, we make the first attempt to tackle this task from a novel game perspective, which effectively learns the uncertain relationship between each vision-language pair with diverse granularity and flexible combination for multi-level cross-modal interaction.Specifically, we creatively model each video frame and query word as game players with multivariate cooperative game theory to learn their contribution to the cross-modal similarity score. By quantifying the trend of frame-word cooperation within a coalition via the game-theoretic interaction, we are able to value all uncertain but possible correspondence between frames and words. Finally, instead of using moment proposals, we utilize the learned query-guided frame-wise scores for better moment localization.Experiments show that our method achieves superior performance on both Charades-STA and ActivityNet Caption datasets.

视频定位博弈论弱监督跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。