通过强化学习提升视频时序定位精度,显著优于现有方法。
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
- 用自适应注意力分配优化模型对视频关键帧的关注
- 在两个测试集上均比当前最佳方法高3.5%准确率
- 适合需要精准时序理解的视频分析任务
时序视频定位(TVG)要求根据语言查询从视频中精确定位相关片段,是视频理解中的难题。视频信息量大且冗余,模型需全面理解全视频以准确检索。我们提出Tempo-R0:一种基于多模态时序感知强化学习的视频多模态大模型。预处理阶段采用基于帧内容变化的自适应注意力分配(SAA),高效利用模型有限注意力;引入显式时间戳-模态对齐(ETA)方法,增强对事件边界的感知能力。微调阶段创新性地采用基于部分无关拒绝的组相对策略优化(PIR-GRPO),使模型不仅接受相关视频-查询对,也拒绝无关对,提升时序推理能力。实验表明,该方法在原始QVHighlights测试集及其修正版(标注更合理)上分别较SOTA提升约3.5%。
原文摘要 · Abstract (English)
Temporal Video Grounding (TVG), which requires pinpointing relevant temporal segments from video based on language query, has always been a highly challenging task in the field of video understanding. Videos often have a larger volume of information and redundancy than texts or images. Models should present comprehensive understanding of the whole video to accurately retrieve query-relevant clips. We thus propose Tempo-R0: a Video Multimodal Large Language Model (Video-MLLM) for the temporal video grounding task via multimodal temporal sensing reinforcement. Specifically, during the preprocessing stage of our pipeline, we employ Self-adaptive Attention Allocation (SAA) method based on frame content variation to efficiently use the MLLM's limited attention. The Explicit Timestamp-modal Aligned (ETA) method is also utilized to strengthen our model's capability to perceive the boundaries of events in the video. In the fine-tuning part of our pipeline, we creatively apply Partial Irrelevance Refusing-based Group Relative Policy Optimization (PIR-GRPO) in TVG area to foster model's temporal reasoning from not only accepting relevant video-query pairs but also refusing irrelevant ones. Experiments demonstrate that our method accomplishes a notable advantage over SOTA solutions by around 3.5% on both the original QVHighlights testbench and its corrected version with more reasonable ground truth annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。