arXiv:2508.07388cs.AI2025-08被引 3

通过逆向任务保持动作理解,提升视频定位精度。

Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability

  • 设计三种逆向任务强化模型动作理解能力
  • 在Charades-STA数据集上提升7.1%定位准确率
  • 适合关注动作语义与定位融合的研究者

时间视频定位(TVG)旨在定位与给定文本查询对应的视频片段,而这些查询通常描述人类动作。然而,我们发现当前方法虽优化高时间交并比(IoU),却常难以准确识别或理解视频与查询中的底层动作,降低有效性。为此,我们提出一种新框架,将基于反演的TVG作为辅助目标,以保持模型的动作理解能力。从原始TVG标注中衍生出三种反演任务:(1) 动词补全,根据视频片段预测查询中被遮蔽的动词;(2) 动作识别,识别查询描述的动作;(3) 视频描述,根据视频片段生成包含查询相关动作的描述。这些任务完全源自原始TVG标注,并在强化学习框架中概率性地与原任务结合。通过精心设计的奖励函数,模型维持了动作理解能力,从而提升定位准确性。实验表明,该方法优于现有最优模型,在3B参数量模型上于Charades-STA数据集上实现[email protected]提升7.1%。

原文摘要 · Abstract (English)

Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods, usually optimizing for high temporal Intersection-over-Union (IoU), frequently struggle to accurately recognize or understand the underlying actions in both the video and query, thus reducing the effectiveness of these methods. To address this, we propose a novel TVG framework that integrates inversion-based TVG as auxiliary objectives to maintain the model's action understanding ability. We introduce three kinds of inversion TVG tasks derived from the original TVG annotations: (1) Verb Completion, predicting masked verbs (actions) in queries given video segments; (2) Action Recognition, identifying query-described actions; and (3) Video Description, generating descriptions containing query-relevant actions given video segments. These inversion tasks are entirely derived from the original TVG tasks and are probabilistically integrated with them within a reinforcement learning framework. By leveraging carefully designed reward functions, the model preserves its ability to understand actions, thereby improving the accuracy of temporal grounding. Experiments show our method outperforms state-of-the-art approaches, achieving a 7.1\% improvement in [email protected] on Charades-STA for a 3B model.

视频定位动作理解逆向任务强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。