用强化学习提升视频理解能力,让AI更懂时间与空间关系。
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

- 通过规则奖励机制优化多模态模型的时空感知能力
- 在时间定位和目标追踪任务上分别提升31.8和31.2分
- 兼顾对话能力,适合构建真实场景下的视频问答系统
强化学习(RL)有助于大语言模型进行复杂推理。受此启发,本文探索将时空特定奖励引入多模态大语言模型(MLLM),以应对视频理解中长时序关联等独特挑战。研究聚焦于基于规则的奖励机制,尤其是时间类奖励,评估其对视频推理的提升效果及泛化能力。提出一种数据高效的强化微调(RFT)方法,在不损失原有能力的前提下,增强特定任务上的视频推理性能。通过在多个时空感知任务上联合微调,构建了VideoChat-R1这一强大的视频多模态大模型。该模型在时间定位(+31.8)和目标追踪(+31.2)等任务上达到当前最优表现,同时提升通用问答基准性能。增强的感知能力与保留的对话能力共同支撑起更可靠的视频对话系统,形成‘时间线索驱动推理’的推断范式。本工作为构建鲁棒、可落地的视频理解智能体奠定基础。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) benefits Large Language Models (LLMs) for complex reasoning. Inspired by this, we explore integrating spatio-temporal specific rewards into Multimodal Large Language Models (MLLMs) to address the unique challenges of video understanding, such as long-range temporal associations. This paper investigates how rule-based rewards, particularly temporal ones, can improve video reasoning and their generalizability. Our study proposes Reinforcement Fine-Tuning (RFT) as a data-efficient method to enhance video reasoning on specific tasks without sacrificing original capabilities. Through joint RFT on multiple spatio-temporal perception tasks, we developed VideoChat-R1, a powerful Video MLLM. VideoChat-R1 achieves state-of-the-art spatio-temporal perception, demonstrating significant improvements in tasks like temporal grounding (+31.8) and object tracking (+31.2), while also improving general QA benchmarks. The enhanced perception and preserved chat abilities contribute to a more reliable video dialogue system, leading to our ``Temporal Clue-driven Reasoning" inference schema. This work provides a foundation for developing robust, real-world video comprehension agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。