arXiv:2507.18100cs.CVcs.AI2025-07EMNLP被引 20

用强化学习提升视频时间定位的精准与泛化能力

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning

  • 两阶段训练:先用优质数据微调,再通过难度可控的强化学习优化
  • 在多个基准上超越现有模型,尤其在开放场景表现更优
  • 开源全部数据、模型和代码,适合研究与工业落地

视频时间定位(VTG)旨在给定自然语言查询时定位视频中的相关时间段。尽管大视觉语言模型(LVLM)和指令微调取得进展,现有方法仍存在时间感知不足和泛化能力差的问题。本文提出一种两阶段训练框架,结合监督微调(SFT)与强化学习(RL),以提升模型精度与鲁棒性。首先利用高质量的精选冷启动数据进行SFT初始化,随后通过难度可控的强化学习进一步增强时间定位与推理能力。在多个VTG基准上的实验表明,该方法在挑战性和开放域场景中均显著优于现有模型。我们深入分析了训练策略与数据集构建,强调高质量冷启动数据与难度控制强化学习的重要性。为促进后续研究与产业应用,我们公开所有中间数据集、模型及代码。

原文摘要 · Abstract (English)

Video Temporal Grounding (VTG) aims to localize relevant temporal segments in videos given natural language queries. Despite recent progress with large vision-language models (LVLMs) and instruction-tuning, existing approaches often suffer from limited temporal awareness and poor generalization. In this work, we introduce a two-stage training framework that integrates supervised fine-tuning with reinforcement learning (RL) to improve both the accuracy and robustness of VTG models. Our approach first leverages high-quality curated cold start data for SFT initialization, followed by difficulty-controlled RL to further enhance temporal localization and reasoning abilities. Comprehensive experiments on multiple VTG benchmarks demonstrate that our method consistently outperforms existing models, particularly in challenging and open-domain scenarios. We conduct an in-depth analysis of training strategies and dataset curation, highlighting the importance of both high-quality cold start data and difficulty-controlled RL. To facilitate further research and industrial adoption, we release all intermediate datasets, models, and code to the community.

视频定位强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。