arXiv:2507.04702cs.CVcs.AI2025-07被引 8

通过强化学习提升视频时序定位精度,显著优于现有方法。

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning

  • 用自适应注意力分配优化模型对视频关键帧的关注
  • 在两个测试集上均比当前最佳方法高3.5%准确率
  • 适合需要精准时序理解的视频分析任务

时序视频定位(TVG)要求根据语言查询从视频中精确定位相关片段,是视频理解中的难题。视频信息量大且冗余,模型需全面理解全视频以准确检索。我们提出Tempo-R0:一种基于多模态时序感知强化学习的视频多模态大模型。预处理阶段采用基于帧内容变化的自适应注意力分配(SAA),高效利用模型有限注意力;引入显式时间戳-模态对齐(ETA)方法,增强对事件边界的感知能力。微调阶段创新性地采用基于部分无关拒绝的组相对策略优化(PIR-GRPO),使模型不仅接受相关视频-查询对,也拒绝无关对,提升时序推理能力。实验表明,该方法在原始QVHighlights测试集及其修正版(标注更合理)上分别较SOTA提升约3.5%。

原文摘要 · Abstract (English)

Temporal Video Grounding (TVG), which requires pinpointing relevant temporal segments from video based on language query, has always been a highly challenging task in the field of video understanding. Videos often have a larger volume of information and redundancy than texts or images. Models should present comprehensive understanding of the whole video to accurately retrieve query-relevant clips. We thus propose Tempo-R0: a Video Multimodal Large Language Model (Video-MLLM) for the temporal video grounding task via multimodal temporal sensing reinforcement. Specifically, during the preprocessing stage of our pipeline, we employ Self-adaptive Attention Allocation (SAA) method based on frame content variation to efficiently use the MLLM's limited attention. The Explicit Timestamp-modal Aligned (ETA) method is also utilized to strengthen our model's capability to perceive the boundaries of events in the video. In the fine-tuning part of our pipeline, we creatively apply Partial Irrelevance Refusing-based Group Relative Policy Optimization (PIR-GRPO) in TVG area to foster model's temporal reasoning from not only accepting relevant video-query pairs but also refusing irrelevant ones. Experiments demonstrate that our method accomplishes a notable advantage over SOTA solutions by around 3.5% on both the original QVHighlights testbench and its corrected version with more reasonable ground truth annotations.

时序定位多模态大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。