用强化学习提升视觉语言模型的视频定位泛化能力
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
- 通过可验证奖励的强化学习框架,引导模型推理视频时间片段
- 仅用2.5K数据即达顶尖性能,且在困难样本上持续优化
- 适合研究视频理解与模型泛化能力的学者使用
时序视频定位(TVG)旨在根据自然语言查询定位长视频中的特定片段,是长视频理解的核心挑战。尽管近期大视觉语言模型(LVLMs)通过监督微调已初显成效,但其泛化能力仍有限。为此,我们提出一种新型后训练框架,通过强化学习(RL)增强LVLM在TVG任务上的泛化能力。主要贡献包括:(1)Time-R1:引入基于推理引导的强化学习后训练框架,采用可验证奖励机制提升模型性能;(2)TimeRFT:在自建的适配强化学习的数据集上探索数据高效策略,使模型逐步掌握复杂样本,提升泛化能力;(3)TVGBench:构建一个精炼但全面的评估基准,涵盖11类查询,视频与查询分布均衡。大量实验表明,Time-R1仅用2.5K训练数据即在多个下游数据集上达到当前最优表现,同时显著提升对长视频的理解能力。
原文摘要 · Abstract (English)
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their abilities to generalize remain limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL). Specifically, our contributions span three key directions: (1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance the capabilities of LVLMs on the TVG task. (2) TimeRFT: we explore data-efficient post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend difficult samples, leading to better generalization. (3) TVGBench: we carefully construct a small yet comprehensive benchmark for LVLM evaluation, assessing 11 types of queries and featuring balanced distributions across both videos and queries. Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using only 2.5K training data, while improving its general video understanding capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。